About
- I'm Tarun — an AI Engineer based in Jaipur, building production-grade LLM evaluation systems, multi-agent pipelines, and full-stack GenAI applications.
- I work end to end with Python, LangGraph, LangChain and LangSmith — from evaluation methodology and retrieval design through to shipped, cost-governed agents.
- On the infra side I reach for FastAPI, Qdrant, Redis, Docker and PostgreSQL, with a bias toward measurable evaluation, drift detection and predictable failure handling.
Connect
Experience
Z
Data Science and AI Intern · ZIDIO Development
Built Python automation pipelines integrating LLM APIs across 5+ sources, cutting downstream data errors ~35–40% and eliminating manual workflows. Deployed prompt-engineering workflows and Power BI dashboards surfacing real-time KPIs, saving ~4–5 hrs/week.
L
Generative AI & Machine Learning Intern · Linux World Pvt. Ltd.
Built ML data pipelines and applied generative AI across two projects — an LLM-based NLP chatbot and a regression-based price predictor. Handled preprocessing and feature engineering for model training and evaluation.
E
B.Tech — Artificial Intelligence & Data Science
Projects
Catching model regressions
in under 4 minutes
LLM Evaluation Harness & Red-Teaming Framework
Personal Project
Benchmarked GPT-4o and Claude Sonnet on 50 MMLU prompts (ROUGE-L, BERTScore, bootstrapped 95% CIs), cutting regression detection from days to under 4 minutes per CI run. LLM-as-judge ensemble with inter-rater agreement scoring — Cohen's κ = 0.81 across 20+ adversarial attack patterns.
PythonLangGraphLangSmithMLflowFastAPIDockerRedis
A RAG pipeline that watches
its own drift
Real-Time RAG Ops Platform with Drift Detection
Personal Project
Production RAG pipeline with Qdrant vector search and cosine-similarity drift detection triggering automated re-indexing — 75% Redis cache hit rate on repeated queries. Real-time ops dashboard tracking p95/p99 latency SLOs, Hit@5 retrieval quality, and embedding drift.
PythonFastAPIQdrantLangChainRedisDockerChart.js
Three agents review your
code before you commit
Multi-Agent Code Review & Auto-Remediation System
Personal Project
3-agent LangGraph pipeline (static analysis, OWASP scanner, LLM fix proposal) running under $0.50/review — found 23 issues including 4 critical vulnerabilities in benchmark testing. 12-pattern OWASP scanner (regex + AST, zero API cost) with typed state contracts.
PythonLangGraphOpenRouterFastAPIDockerPostgreSQLRedis
Code explained out loud,
in English and Hindi
CodeCave — Bilingual LLM Code Tutor
Personal Project
Bilingual (English/Hindi) LLM code tutor that turns pasted Python/JavaScript into spoken step-by-step explanations, using a strict JSON output schema, server-side validation, and prompt-injection mitigation so malformed output never reaches the UI. Provider-fallback layer (Gemini models → Groq) with runtime model discovery and 503/429 retries, plus response caching and a rule-based offline explainer for free-tier resilience; API keys stay server-side.
Gemini APIGroq APINode.js (Vercel Serverless)JavaScriptWeb Speech APIWeb Workers
Knows when not to act alone
HITL Approval Agent — Stateful AI Safety System
Personal Project
Stateful human-in-the-loop approval agent that gates execution on a self-assessed confidence score — auto-executing above threshold, else persisting the task with a full JSON audit trail for human approve/reject. Surfaced an architectural safety gap motivating per-category approval thresholds.
Next.js 14PrismaPostgreSQLGroq Llama 3.3-70B
Skills
Python
SQL / PostgreSQL
JavaScript / TypeScript
LangChain
LangGraph
LangSmith
Qdrant
FastAPI
Docker
Redis
Prisma
Next.js
MLflow
Groq API
OpenRouter
Git
GitHub Activity
Loading contribution graph…
Achievements
Generative AI
Mastermind Outskill
Data Science & Machine Learning
Udemy
Data Analytics
Google
LLM & GenAI Systems
RAG & Retrieval
Agentic Evaluation
Infra & Deployment
Still reading? That means something clicked.
Let's talk.
Send an email
Let's talk.