Tarun Singh Chauhan
Tarun Singh Chauhan
Builds agents that don't hallucinate (mostly)
πŸ“ Jaipur, IN
--:--:--

I'm an AI Engineer based in Jaipur, and I build systems that hold up after the demo ends β€” LLM evaluation harnesses, multi-agent pipelines, and RAG infrastructure that don't fall over the first time real traffic hits them.

Most days that's 🐍 Python and 🦜 LangChain with πŸ•ΈοΈ LangGraph, wired up to πŸ” Qdrant for retrieval and πŸ“Š LangSmith for tracing. Lately a lot of that means benchmarking GPT-4o against Claude Sonnet, running LLM-as-judge ensembles, and building red-team suites that try to break my own agents before a user does.

Two internships and five-plus shipped projects in, the part I still like most is simple: something is unreliable, and by the time I'm done, it isn't.

Story So Far
2024 β€” Present
Z
ZIDIO Development
Jun – Aug 2025
AI Engineer Intern Β· Remote
  • Built AI-powered Python pipelines integrating LLM APIs to extract, clean, and process structured data from 5+ sources, cutting downstream data errors by ~35–40% and killing off manual workflows.
  • Shipped prompt-engineering workflows and automated reporting tools in Python and SQL, plus Power BI dashboards surfacing real-time KPIs β€” saving ~4–5 hours every weekly cycle.
  • Worked Git-based version control and structured code review into the team's day-to-day, keeping production automation scripts reliable and consistent.
  • Wrote pytest suites and added structured logging across the pipelines so failures got caught before they reached reporting.
L
Linux World Pvt. Ltd.
Jul – Aug 2024
Generative AI & Machine Learning Intern Β· Jaipur
  • Built ML data pipelines and applied generative AI across two projects β€” an NLP chatbot on LLM APIs and a regression-based price predictor β€” producing production-ready training datasets.
  • Cleaned and structured raw datasets with pandas and NumPy for model training and evaluation.
  • Paired with the ML team on experimentation and prompt design, documenting results as approaches improved.
Tools I Use

Python underneath almost everything, LangGraph for anything that needs to reason in steps, Qdrant when it needs to remember, and Docker so it works the same on my machine as anywhere else.

🐍 Python 3.12 🐘 PostgreSQL ⚑ FastAPI 🐳 Docker πŸŸ₯ Redis 🦜 LangChain πŸ•ΈοΈ LangGraph πŸ“Š LangSmith πŸ” Qdrant ⚑ Groq πŸ”€ OpenRouter πŸ”§ Git & GitHub Actions βœ… pytest πŸ§ͺ pydantic πŸ“ˆ MLflow
Things I Have Built
github.com/TarunSinghChauhan

Cinematic, line-by-line
AI code walkthroughs

CodePulseshipped
Personal Project

Sends code to LLaMA 4 via Groq and gets back a structured execution script that powers a live, animated walkthrough β€” variable tracking, call stack, voice narration, and a "Break Mode" that injects bugs on purpose and narrates the fix.

Next.js 14TypeScriptGroq Β· LLaMA 4Framer MotionWeb Speech API

Catching model regressions
in under 4 minutes

LLM Evaluation Harness & Red-Teaming Frameworkshipped
Personal Project

Benchmarks GPT-4o against Claude Sonnet across 50 MMLU prompts with ROUGE-L, BERTScore, and bootstrapped confidence intervals. An LLM-as-judge ensemble (Cohen's ΞΊ = 0.81) plus 20+ adversarial red-team patterns catches safety regressions before they ship.

LangGraphLangSmithMLflowFastAPIDocker

Three agents review your
code before you commit

Multi-Agent Code Review & Auto-Remediationshipped
Personal Project

A 3-agent LangGraph pipeline β€” static analysis, OWASP security scan, LLM fix proposal β€” under an enforced $0.50 per-review budget. Caught 23 issues in benchmark testing, 4 of them critical: SQL injection, hardcoded secrets, insecure deserialization, shell injection.

LangGraphOpenRouterFastAPIPostgreSQLRedis

A RAG pipeline that watches
its own drift

Real-Time RAG Ops Platform with Drift Detectionshipped
Personal Project

Qdrant vector search with zero-API-cost embeddings and cosine-similarity drift detection that triggers automatic re-indexing. A live ops dashboard tracks p95/p99 latency, Hit@5 retrieval quality, and drift score β€” 75% Redis cache hit rate on repeat queries.

QdrantLangChainRedisChart.jsDocker

Full company research reports, under 60 seconds

Autonomous Financial Research Agentshipped
Personal Project

A tool-augmented agent that runs a 4-step reasoning chain β€” market context, financial analysis, risk assessment, investment thesis β€” over Yahoo Finance and DuckDuckGo, with a per-query cost budget of $0.50. Every report ships with a SHA-256 reproducibility hash and a full tool-call audit trail; Redis caching cuts repeated-query cost by ~60%.

OpenRouterYahoo Finance APILangSmithRedisDocker

Let's talk

Building something that needs an agent, an eval harness, or a RAG pipeline that doesn't drift silently? Email is the fastest way to reach me.

Email me
tarunsinghchauhan088@gmail.com Β· +91 7877586779
counting visitors…

Where the work travels

rotating live Β· local time per city

⟲ spinning west β†’ east, same direction as Earth