Open to new roles
  • New York

Hi there, nice to meet you.

Rishi Chhabra

I'm an AI engineer in New York. I build LLM platforms and RAG systems, and the checks that tell you whether they actually work.

01 About me

Portrait of Rishi Chhabra
New York, NY

Right now I'm at Ikka Labs, building the internal AI tools other engineers lean on. The biggest one is Apex, a self-service platform that 4+ teams use to spin up 13 different data and inference engines on their own.

I also built PRaaS, a code-review agent teams can switch on for their repos. These days it reads about 117 pull requests a week.

Before Ikka, I spent my last two semesters at AriesView turning keyword search over 3,000 documents into a proper RAG system, and context precision went from 0.61 to 0.94. I finished my master's in Machine Learning at Stevens in May 2026. My first job was in Delhi, building mobile apps for a startup.

Currently
AI Engineer at Ikka Labs
Based in
New York
Studied
M.S. in Machine Learning, Stevens
Looking for
AI engineering, AI platform, MLOps, or forward deployed roles

Highlights

4+

teams use Apex, the platform I built

Ikka Labs
117

pull requests my review agent reads each week

Ikka Labs
0.61 → 0.94

context precision after I rebuilt retrieval

AriesView
34%

faster API responses after I swapped Redis for Valkey

AriesView
02 Work

Where I've shipped.

New York, NY
  • I built Apex, a self-service platform that 4+ teams use to provision 13 data and inference engines across on-prem Kubernetes, Amazon EKS, and Azure AKS.
  • I made those deployments safe to repeat, with policy checks, signed previews before anything goes out, secrets kept in HashiCorp Vault, and an AI-guided setup flow.
  • I created PRaaS, a code-review agent that runs in GitHub Actions. Teams opt in, and it now reviews about 117 pull requests a week across 13 repositories, picking between seven Gemini, Amazon Bedrock, and self-hosted models.
Ikka Labs · 2026
Boston, MA (Remote)
  • I moved a 3,000-document enterprise corpus off keyword-only search and onto production RAG, with semantic chunking, hybrid BM25 and semantic retrieval in Weaviate, GraphRAG on Neo4j, and Cohere reranking. On a 100-query test set, context precision went from 0.61 to 0.94.
  • I set up the evaluation pipeline with RAGAS, DeepEval, custom metrics, and LLM-as-a-judge, and put it in CI so quality regressions show up in a pull request instead of in production.
  • I shipped a B2B chatbot on top of that retrieval stack and worked directly with the business clients until it was delivered.
  • I swapped Redis for Valkey on our self-hosted caching, and the API endpoints got 34% faster.
AriesView · 2025
Delhi, India
  • I built three cross-platform apps with Flutter, Node.js, MongoDB, and S3: iManager for football teams, Panda for dating, and Paradisseo for events and meeting people. All three launched as MVPs on the App Store and Google Play.
  • I added JWT login, Apple Pay and Stripe payments, and push notifications, and worked with the founders and their clients the whole way through.
  • I moved the infrastructure from Akamai Linode to AWS with Terraform (EC2, S3, Route 53, IAM, Elastic Beanstalk, ElastiCache, load balancing) and wired in Amazon Bedrock.
Incuwise · 2024
03 Projects

Built on my own time, measured anyway.

01

Cost-Aware Semantic Gateway for LLM Routing

Sending every question to GPT-4o is expensive, and most questions don't need it. So I put a small classifier in front that guesses how hard each one is. Easy ones stay on Llama 3.2, running locally on vLLM, and only the hard ones go to GPT-4o.

68.1%lower cloud-model spend
79%of queries kept local
93.4%routing accuracy

The spend and share numbers come from an 81-query benchmark. Routing accuracy was measured on labeled prompts.

FastAPI · vLLM · Llama 3.2 · GPT-4o · Streamlit · Docker Source ↗
02

Automated LLM Evaluation & Observability Pipeline

I wanted a bad RAG change to fail the build, the same way a broken test does. Every pull request gets scored on faithfulness, answer relevancy, and contextual precision, and if any score drops below 95% the merge is blocked. With LangSmith tracing, working out why a check failed takes minutes instead of hours.

95%floor on every metric
3metrics per pull request
RAGAS · DeepEval · LangSmith · GitHub Actions · Claude Haiku Source ↗

04 Toolbox

Cloud & Containers

  • Kubernetes
  • Docker
  • OpenShift
  • AWS
  • Amazon Bedrock
  • Azure

Infrastructure & Delivery

  • Terraform
  • Kustomize
  • GitHub Actions
  • GitLab CI/CD
  • MLflow
  • Kubeflow
  • OAuth 2.0/OIDC
  • RBAC
  • HashiCorp Vault

AI Inference & Serving

  • vLLM
  • llm-d
  • Ollama
  • GPT-4o
  • Llama
  • Gemini
  • PyTorch
  • Hugging Face

Retrieval & Evaluation

  • RAG
  • GraphRAG
  • Hybrid search
  • BM25
  • Semantic chunking
  • Cohere reranking
  • LangChain
  • LangGraph
  • MCP
  • Tool calling
  • ReAct
  • RAGAS
  • DeepEval
  • LangSmith
  • LLM-as-a-Judge

Backend & Data

  • Python
  • SQL
  • JavaScript
  • Linux
  • FastAPI
  • PostgreSQL
  • Valkey
  • Redis
  • ElastiCache
  • Kafka
  • Weaviate
  • Neo4j

05 Education

Degrees

Stevens Institute of Technology

M.S., Machine Learning · Hoboken, NJ

Sep 2024 to May 2026

Central University of Haryana

B.Tech., Computer Science · India

Sep 2020 to Jun 2024

Certifications

  • AWS Certified Machine Learning Engineer, Associate2025
  • AWS Certified Developer, Associate2025
  • Databricks Certified Generative AI Engineer2026
  • Databricks Certified Associate Developer for Apache Spark2025
  • Anthropic Claude Certified Architect, Foundations2026
06 Contact

Want to work together? Say hello.

I'm looking for a full-time role in AI engineering, AI platform, MLOps, or forward deployed engineering. I'm based in New York and open to relocating. Email is the quickest way to reach me.

Rishi.Chhabra@outlook.com