Open to full stack AI roles · Delhi, India / Remote

I build AI systems.
End to end.

Harshit Rai — Full Stack AI Engineer

I design and ship complete AI systems — classic ML, LLM/RAG pipelines, and agentic automation workflows — on a backend I build myself in Python, Go, Rust, and TypeScript, deployed with real Kubernetes/DevOps, not notebooks. Currently working as an AI Quality Engineer, where I also stress test models before they reach real users.

17+
systems shipped
4
backend languages
8
domains mastered
30–40%
review time cut
01
Full stack, not just the model
The LLM is the easy 10%. I write the Go policy gate, the Rust training core, the TypeScript frontend, and the Kubernetes it all runs on — because a model nobody can deploy isn't shipped, it's a notebook.
02
Agentic, with guardrails
Every multi agent system I ship is schema validated and policy gated — an agent's opinion is never a thing that executes on its own. Reasoning and authority stay separated on purpose.
03
Proof, not claims
44.4% → 94.4% accuracy, verified on held out rows. 95% fewer parsing timeouts, traced to a real memory leak. Numbers I can show you, not a slide that says "AI powered."
// profile
A full stack AI engineer
Not a prompt and pray builder — I own the model, the backend, and the infra it runs on.

I'm a Full Stack AI Engineer based in Delhi, India, currently working as an AI Quality Engineer at Smarter.Codes. My range runs the full width of the stack: classic ML, fine tuned LLMs and RAG, multi agent and agentic automation workflows — all sitting on a backend I write myself in Python, Go, Rust, and TypeScript, deployed as real microservices on Kubernetes with actual CI/CD, not a demo notebook.

The lens I bring to all of it is quality first — golden datasets, error taxonomies, and benchmark harnesses that tell you whether a system is actually good, not just whether it runs. Always looking for harder problems at the intersection of AI, automation, and infrastructure.

locationDelhi, India
emaildev.harshitrai@gmail.com
githubgithub.com/Harshitraiii2005
focusAgentic AI · RAG/LLM · Go/Rust/TS Backend · DevOps Automation
system.status
// live profile
Current roleAI Quality Engineer
OrgSmarter.Codes
Backend stackPY · GO · RUST · TS
Eval coverage85–90%
Systems shipped16+
Model checksPASSING
// 02 · experience
Where I've shipped
Four roles, ~1.5 years, one thread: building AI systems end to end and proving they hold up.
role.active
AI Quality Engineer
FULL TIME
Smarter.Codes · Panchkula, Haryana, India · On site · Jul 2026 – Present · 3 mos
  • Engineered end to end LLM evaluation pipelines in Python scoring accuracy, relevance, groundedness, and consistency against custom quality rubrics across 50+ production test prompts.
  • Built a benchmarking framework on 100+ golden reference queries to continuously track model performance and catch prompt/response regressions before they shipped.
  • Stress tested RAG pipelines end to end — retrieval relevance, ranking quality, answer grounding — surfacing 15–20 critical edge cases (hallucinations, irrelevant retrievals, incomplete context) that would otherwise have reached users undetected.
  • Engineered accuracy, latency, and indexing improvements for an AI medical writer system, optimizing retrieval/indexing state and the response pipeline — a 30% performance improvement in output quality and response time.
  • Built automated eval harnesses and Python test scripts to scale LLM output scoring, cutting manual review effort by 30–40%.
  • Owned functional and regression testing across web applications, APIs, and AI driven systems, sustaining ~85–90% test coverage on core modules.
  • Engineered validation for structured and unstructured document parsing (PDF, DOCX, tables) feeding production AI pipelines, integrated into a mini CI/CD setup.
role.freelance
Machine Learning Evaluator
FREELANCE
Handshake · Noida, Uttar Pradesh, India · Remote · Dec 2025 – May 2026 · 6 mos
  • Worked as an AI Trainer with Handshake AI, partnering with leading AI research labs to evaluate and improve large language model outputs.
  • Designed and executed model evaluations across reasoning, coding, and factual accuracy tasks — identifying failure modes and delivering structured feedback that fed fine tuning and RLHF pipelines.
  • Benchmarked model performance across versions to surface edge cases and drive measurable gains in reliability, alignment, and overall output quality.
role.self directed
Optimization Specialist
SELF DIRECTED
IntelliAssist · Oct 2025 – Nov 2025 · 2 mos
  • Integrated a LoRA fine tuned, locally hosted LLM into a RAG pipeline with FAISS vector search and Sentence Transformers embeddings.
  • Applied token usage optimization (95% parameter reduction) and prompt engineering to cut hallucination by 60% and lift domain accuracy from 35% to 75%.
  • Implemented latency and cost aware inference design, hitting sub 200ms response times on CPU only infrastructure, with logged generation quality metrics for continuous evaluation.
role.internship
Machine Learning Intern
INTERNSHIP
Brainwave Matrix Solutions · Noida, Uttar Pradesh, India · Remote · Jul 2025 – Sep 2025 · 3 mos
  • Built end to end MLOps pipelines deploying NLP models (BERT, mDeBERTa) to Kubernetes with HPA auto scaling and Ingress routing — cutting model deployment time by ~60%.
  • Designed a 9 stage DevSecOps CI/CD pipeline in Jenkins integrating SonarQube (SAST), Trivy (CVE scanning), and OWASP dependency checks — blocking HIGH/CRITICAL vulnerabilities before production.
  • Containerized FastAPI ML services with Docker, optimizing image size from 1.4GB to 148MB (90% reduction) through multi stage builds and layer caching.
  • Implemented a GitOps deployment workflow with ArgoCD, enabling zero downtime Kubernetes rollouts from GitHub commits via automated manifest updates.
  • Developed and deployed end to end ML models using Python and TensorFlow with modern MLOps practices — achieved ~45% performance gain via pruning and quantization, with automated GitHub Actions CI/CD for testing, versioning, and release.
// 03 · projects
Systems I've built
17 projects across agentic AI, RAG/LLM, classic ML, and the Go/Rust/TypeScript backends & DevOps that ship them — all live, not just prototyped.
flagship.01
TREEHASH LIVE
treehash — Vectorless RAG Engine
Python · SQLite (FTS5) · PostgreSQL · BM25 · FastAPI · MCP · Docker
  • Parses documents into a hashed address tree and resolves most queries by direct lookup — no embeddings, no vector index, O(1) average latency.
  • Falls back to BM25 full text search only when structural resolution misses, with an optional semantic tier gated behind both failing first.
  • Cut parsing/indexing stage timeouts by 95% by tracing and fixing a page cache memory leak in the PDF parser, holding ingestion stable at 100MB+/multi document scale.
  • Hash chained audit trail and real provider billed token accounting per query; published to PyPI as treehash-rag.
Zero Vector DB92% Accuracy95% Fewer TimeoutsPyPI Published
flagship.02
AP AUTOPILOT LIVE
AP Autopilot — Accounts Payable Automation
Temporal · Go · PostgreSQL · Next.js · MCP · QuickBooks/Xero/NetSuite/Stripe
  • Separates reasoning from authority — LLMs extract invoice data only; a deterministic Go policy gate makes every payment decision.
  • 94.2% field extraction accuracy with a 0% false approval rate; auto approves low risk invoices in ~0.5s, escalates the rest to named approvers.
  • Hash chained, replayable audit ledger on durable Temporal workflows that survive restarts without losing state.
0% False Approval94.2% ExtractionBounded Autonomy
flagship.03
VERIFACT AI LIVE
VeriFact AI — Zero Hallucination RAG
Hybrid Retrieval · Multi Tenant Isolation · Encryption at Rest · MIT Core
  • Answers strictly from the user's own documents — a verbatim quote or an honest "I don't know," never an invented answer.
  • Confidence scored hybrid retrieval benchmarked on 94 real questions, with per user isolation and encryption at rest for multi tenant use.
  • Modular, MIT licensed retrieval core kept cleanly separable from the product shell.
Zero Hallucination94 Q BenchmarkMulti Tenant
flagship.04
SENTRYOPS LIVE
Sentryops — Autonomous DevOps Agent
Rust · Go SDK · Kubernetes · OPA style Policy · Groq · MCP
  • Discovers Kubernetes pods, runs multi endpoint health checks, and auto files tickets the moment a service degrades.
  • Clusters raw logs into causal dependency graphs, then an LLM diagnoses root cause and proposes remediation — ranked by actual cause, not symptom.
  • Auto resolves known safe errors behind an OPA style default deny policy gate, with correlation ID audit trails on every decision.
Rust CoreCausal RCADefault Deny Gate
flagship.05
LEGAL CONTRACT PIPELINE LIVE
Legal Contract Pipeline
LangGraph · FastAPI · Pinecone · PostgreSQL · Celery/Redis · React · Jenkins · MLflow · Kubernetes
  • Automated first pass legal review via a 5 agent workflow (Extractor, Scorer, Compliance Checker, Redliner, Writer).
  • Grounded LLM risk assessment in precedent — RAG with Pinecone retrieves top 3 similar historical clauses.
  • Gated deployments on a 60%+ calibration accuracy threshold via a Jenkins + MLflow eval harness.
5 Agent Workflow60%+ Eval GateRAG Top 3
flagship.06
DEVOPS HELPDESK AGENT LIVE
Autonomous DevOps Helpdesk Agent
Python · Agentic Workflows · FastAPI · SQLite · Vector Store · Kubernetes · Jenkins
  • Cut incident response time with an autonomous agent that plans, calls tools, and reasons over runbooks and live alerts.
  • Tiered safety guard system — auto approves read only calls, routes state mutating actions to manual sign off.
  • Full audit traceability via an immutable, append only Postgres audit log.
Human in the LoopParallel ExecImmutable Audit Log
flagship.07
DEVINTEL LIVE
DevIntel — Competitive Intelligence Pipeline
FastAPI · Next.js · PostgreSQL · Redis · Playwright · Claude/GPT 4o/Groq · Docker
  • Full stack SaaS that monitors competitor URLs on a schedule, detects changes via SHA 256 hashing, and generates LLM intelligence reports.
  • Multi channel report delivery (Slack, Email, Webhook) via Redis/RQ async jobs and APScheduler cron.
  • Estimated 3–5 hours/week saved per engineering team through end to end automation.
SHA 256 Diffing3–5 hrs/wk savedMulti channel
flagship.08
INTELLIASSIST RESEARCH
IntelliAssist — RAG Enhanced Virtual Assistant
PyTorch · Hugging Face · FAISS · LoRA · Whisper · Gradio · Docker
  • 3.4× context relevance over a vanilla RAG baseline via hybrid FAISS + BM25 search, MMR ranking, and cross encoder reranking.
  • LoRA adapters on GPT 2 cut fine tuning params by 95%, hitting a 96% pass rate across 25 eval cases.
  • Sub 200ms average response latency with a multimodal interface and hallucination detection fallback.
3.4× Relevance96% Pass Rate95% Param Cut
flagship.09
MODELFORGE LIVE
ModelForge — Agentic Model Remediation
Express/TS · Go · Rust · MongoDB/PostgreSQL/S3 · OpenAI/Claude/Gemini (BYOK)
  • Multi agent pipeline (QA → Diagnose → Research → Strategize) reasons over real computed metrics — every agent response is schema validated before Go dispatches it; no agent ever emits code that runs.
  • Trains real remediation candidates in Rust (reweighing, resampling, hyperparameter search, neural net, RAG grounding) and picks a winner via Pareto frontier selection on identical held out rows.
  • Verified run: 44.4% → 94.4% accuracy (+50pp) on an 18 row held out split — same rows scored before and after, no cherry picking, no mocked numbers.
Real Rust Retrains+50pp VerifiedSchema Guardrailed Agents
// all systems
01 · FEATURED
IntelliAssist — RAG AI Assistant
Production RAG system on fine tuned GPT 2 + LoRA. 96% pass rate across 25 test cases, 0.15s avg response, hybrid FAISS+BM25 search, multimodal I/O.
PythonLoRAFAISSWhisperDocker
PaperIntel AI — Research Paper Analyzer
Multi agent LangGraph platform that summarizes papers, flags methodology issues, and produces PPTX presentations. FastAPI + PostgreSQL + Redis.
LangGraphFastAPIGroqK8s
03 · LIVE
Indic NLP Microservice
Multilingual sentiment (5 star BERT) + zero shot topic classification (mDeBERTa) for English, Hindi, Hinglish. Async batch processing, K8s + HPA.
BERTmDeBERTaK8sAWS
04 · LIVE
HireSparkAI — Talent Intelligence
Dual RAG pipelines for freshers vs experienced hires. Explainable matching, interview question generation, resume optimization. ArgoCD + Jenkins CI/CD.
FAISSLangChainArgoCD
Secure MLOps Recommendation API
Neural Collaborative Filtering engine (MovieLens 100k) with DevSecOps hardening — Bandit, Trivy, Dockle, Cosign signing, MLflow tracking.
PyTorchMLflowTrivyK8s
LangChain Chain Patterns
10 working LangChain chain pattern implementations — sequential, router, memory, tool use — each with docs and graph visualizations.
LangChainGroqPython
07 · LIVE
DocuScribe — AI Document Processor
Extracts structured insights from PDFs/images via Claude vision API, generates SRT subtitles, auto tags docs. Handles 100+ docs/day.
Claude VisionCeleryRedis
TurboLearn — Adaptive Learning
AI personalized learning paths with SM 2 spaced repetition, AI tutoring chat, and 2,000+ concurrent users at sub 200ms latency.
Next.jsFastAPIPostgreSQL
09 · LIVE
DevIntel — Competitive Intelligence
Schedules competitor URL monitoring, SHA 256 diffing, LLM report generation, and multi channel delivery via Slack/email/webhook.
Next.jsPlaywrightRedis
10 · AGENTIC · LIVE
DevOps Helpdesk Autonomous Agent
Parses infra alerts, reasons over conflicts, executes safe actions behind a human approval gate, with full ID correlated audit logs.
FastAPIVector StoreGroq
11 · FEATURED · LIVE
Legal Contract Pipeline
Multi agent legal compliance pipeline via LangGraph, Pinecone risk calibration, and a 60% accuracy eval gate blocking bad merges.
LangGraphPineconeMLflow
12 · LIVE
Python Q&A Assistant
RAG system grounded in 50k Stack Overflow pairs — MiniLM embeddings, ChromaDB retrieval, Groq generation, ~700ms avg latency.
ChromaDBSentence TransformersGroq
13 · FEATURED · LIVE
treehash — Vectorless RAG Engine
Parses documents into a hashed address tree, resolving most queries by O(1) average lookup instead of vector search. Cut parsing/indexing stage timeouts by 95% via a page cache memory leak fix; BM25 fallback, hash chained audit trail, published on PyPI.
PythonSQLite FTS5PostgreSQLMCP
14 · FEATURED · LIVE
AP Autopilot — AP Automation
Invoice extraction, matching, and risk scoring with a deterministic Go policy gate that alone can approve payment. 94.2% extraction accuracy, 0% false approvals.
TemporalGoPostgreSQLNext.js
15 · LIVE
VeriFact AI — Zero Hallucination RAG
Answers strictly from a user's own documents with confidence scored hybrid retrieval — a verbatim quote or an honest "I don't know," benchmarked on 94 real questions.
Hybrid RetrievalMulti TenantMIT Core
16 · AGENTIC · LIVE
Sentryops — Autonomous DevOps Agent
Discovers K8s pods, clusters logs into causal dependency graphs, and diagnoses root cause via LLM, auto resolving known safe errors behind a default deny policy gate.
RustKubernetesOPA PolicyGroq
17 · FEATURED · LIVE
ModelForge — Agentic Model Remediation
Multi agent pipeline diagnoses an underperforming model, trains real Rust remediation candidates, and picks a winner by Pareto frontier selection — verified 44.4% → 94.4% on held out data.
GoRustExpress/TSMulti Agent
// 04 · capability stack
One engineer, the whole stack
Classic ML to agentic automation, Go/Rust/TS backends to DevOps — I ship the full pipeline, not just the model.
01// classic ml
Regression & Classification XGBoost / Random Forest Feature Engineering Clustering & Unsupervised Time Series Forecasting CNN / RNN / Transformers
02// llm & generative ai
LLM Application Design Fine Tuning / LoRA / PEFT Prompt Engineering Context Engineering Claude / GPT 4o / Groq Whisper / Multimodal I/O
03// agentic ai & automation
Multi Agent Orchestration LangGraph / Tool Use Workflow Automation (Temporal) Human in the Loop Design Autonomous Workflows MCP (Model Context Protocol) Policy Gated Execution
04// rag & retrieval
Hybrid Search (Vector + BM25) FAISS / Pinecone / ChromaDB Reranking & MMR Chunking & Structural Retrieval Grounded / Zero Hallucination QA LangChain
05// optimization
Latency & Cost Optimization Quantization & Pruning Token Usage Reduction Parameter Efficient Fine Tuning Inference Scaling
06// evals & testing
Eval Harness Design Golden Dataset Curation Error Taxonomy Analysis Regression Testing Benchmarking LLM as Judge
07// backend
Python FastAPI Go Rust TypeScript Microservices PostgreSQL Celery / Redis Async / REST APIs
08// devops & mlops
Kubernetes / Docker Jenkins / GitHub Actions ArgoCD / GitOps MLflow / DVC AWS / Terraform
// 05 · faq
Questions people ask
The things recruiters and hiring managers actually want to know.
Are you a model person or a backend person?+

Both. I fine tune and evaluate models, but I also write the Go/Rust services that decide, the TypeScript that renders, and the Kubernetes that runs it — a model nobody can deploy isn't shipped.

Are the metrics on this site real?+

Yes — every number here (44.4%→94.4% accuracy, 95% fewer timeouts, 0% false approval rate) comes from a verified run or a live deployment, not a marketing estimate.

What's your actual day to day stack?+

Python and FastAPI for ML/eval work, Go and Rust where determinism or performance matters, TypeScript/Next.js for the frontend, and Kubernetes/Docker/Jenkins/ArgoCD to ship all of it.

What are you looking for next?+

Full stack AI engineering roles — building and shipping agentic or RAG systems end to end, not just prompting one. Open to full time roles and hard technical collaborations.

// 06 · contact
Open a channel
Best for collaboration, roles, or hard full stack AI problems.
LOCATION
Delhi, India