NKNantha Kumar
Get in touch
Open to AI Engineer & Generative AI roles

Nantha Kumar

AI Engineer · Software Engineer Irving, TX · open to relocation

I build production LLM systems — agents, RAG, and evals — that hold up outside a demo. Generative AI on LangGraph with grounded retrieval and multi-provider routing, full-stack on Next.js and FastAPI, and open-source work merged into Microsoft PyRIT.

Nantha Kumar
Nantha Kumar AI Engineer
Irving, TX Open to relocation
4
LLM projects built end to end
3
Open-source PRs to Microsoft & Keras
1,014
Documents indexed for grounded retrieval
720
Core samples modeled in ML research
Selected work

Things I've built 4 projects

LLM systems and applied ML — built end to end, from retrieval and routing to auth, billing and deploy.

Agentic AI platform with voice and memory

  • Built a six-stage LangGraph agent with provider-agnostic routing across Claude, Gemini 2.0 Flash and Llama 3.3 70B — switchable through a single environment variable — holding cost near $0.0002 per session.
  • Built RAG memory on Postgres with pgvector and HNSW indexing (768-dim embeddings, custom cosine-similarity SQL) returning results in under 200 ms.
  • Added browser push-to-talk voice via Gemini speech-to-text and Eleven Labs, deployed across Vercel, Hugging Face Spaces and Supabase with LangSmith tracing, Clerk auth and Stripe billing.
Next.jsFastAPILangGraphpgvectorClaudeGeminiStripe

FleetIQ ↗

Buildathon '26

Grounded document Q&A agent — team project · Dallas, 32-hour build

  • Built a hybrid retrieval agent over 1,014 fleet documents spanning 12 real carrier form types that answers with citations and abstains when the evidence is missing.
  • Generated the corpus from 12 templates with realistic entity variation, creating known ground truth so retrieval quality could be measured rather than eyeballed.
  • Delivered backend, dispatch interface and pitch inside the 32-hour window.
LangGraphFastAPIHybrid RAG

Autonomous job-application pipeline

  • Ingests postings from four job boards, scores them with a heuristic filter followed by an LLM fit review, tailors the resume through a LangGraph workflow and renders the PDF — run end to end against live postings.
  • Added a fact-verification node that checks every generated claim against a structured profile and blocks invented skills, titles or dates before the document is rendered.
PythonLangGraphGreenhouseLeverAshbyWorkable

AI Career Job Advisor

LLM-powered career SaaS — 3-person Agile team

  • Owned resume-to-job-description matching, ATS scoring and cover-letter generation with schema-validated outputs.
  • Wrote 25+ backend tests and contributed to a Selenium end-to-end suite.
OpenAIFastAPINext.jsPostgreSQL
Open source

Contributions upstream 3 PRs

Fixes and features sent to tools the ML community actually runs — merged into Microsoft PyRIT and vasim, in review at Keras.

LLM red-teaming toolkit · PR #2649

  • Added a phrase-aware PinyinConverter that rewrites Chinese prompts into Pinyin variants, testing whether model safety holds when the same request crosses scripts.
  • Shipped with unit and determinism tests, documentation and packaging updates — +570 / −49 across 10 files.
PythonLLM securityPytest

VM autoscaling simulator · PR #139

  • Root-caused an off-by-one in scaling-event counts — a trailing NaN from pandas shift(-1) was counted as a change, skewing the Pareto-frontier ranking.
  • Fixed the count and guarded a KeyError in ParetoFrontier the fix made reachable; added unit tests and corrected four end-to-end expectations — +122 / −8 across 5 files.
PythonpandasPytest

Keras ↗

In review

Deep-learning framework · PR #23639

  • Fixed silent int→float truncation in selu, soft_shrink and sparse_plus — selu was degrading to elu on integer inputs — across the NumPy, TensorFlow, PyTorch and OpenVINO backends and the symbolic output spec.
  • Added tests covering the corrected behavior on every backend.
PythonTensorFlowPyTorch
Research

Applied ML research Aug 2025 — present

AI / Machine Learning Research Assistant

iResearchE³ Lab

University of Texas at Arlington

scikit-learnXGBoostTensorFlowKerasOptunaSHAP
  • Built the lab's end-to-end ML pipeline in Python to predict rock permeability from NMR T2 measurements across 720 heterogeneous carbonate core samples, handling targets spanning six orders of magnitude with log-space modelling and CDF-based feature encoding.
  • Tuned XGBoost, Random Forest and deep neural-network regressors with Optuna (TPE) under repeated 5-fold × 2 cross-validation across three nested feature scenarios, improving test R² by ~50–60% over the lab's SDR physics baseline (R² ≈ 0.42 RF/NN, 0.40 XGBoost vs. 0.26 SDR).
  • Stress-tested robustness across seven training sizes (120–720 samples) and 50 repeated 80/20 splits to map how accuracy scales with data volume.
  • Used SHAP and Spearman rank correlation to show that 4 of 11 features drive roughly 95% of the output. Co-authoring the manuscript.
Toolkit

What I work with

AI & LLM engineering

LangGraphLangChainRAG (hybrid & vector)Prompt engineeringFunction callingStructured outputsLLM evaluationMulti-provider routingGrounding & guardrailsSTT / TTSLangSmithClaudeOpenAIGeminiGroqLlama 3.3

Machine learning

scikit-learnXGBoostRandom ForestTensorFlowKerasPyTorchCNNLSTMOptunaSHAPCross-validationHyperparameter tuningModel interpretabilityNumPypandas

Languages

PythonTypeScriptJavaScriptJavaC++CSQL

Frameworks & web

Next.jsReactFastAPINode.jsExpressRESTGraphQL

Data & cloud

PostgreSQLpgvector (HNSW)SupabaseVector embeddingsDockerVercelHugging Face SpacesGitGitHub ActionsStripeClerk

Testing & tooling

JestPytestSeleniumClaude CodeCursorGitHub Copilot
Background

Education & publication

M.S. Computer Science

University of Texas at Arlington — Arlington, TX

Dual specialization — Software Engineering & Intelligent Systems (AI/ML)

Aug 2024 – May 2026 · GPA 3.60 / 4.00

B.Tech Information Technology

Rajalakshmi Engineering College — Chennai, India

Aug 2020 – May 2024 · CGPA 8.28 / 10.00, First Class

Publication IEEE

Sign Language Caption Generation Using LSTM — IEEE International Conference for Convergence in Technology (I2CT), 2024. Real-time sign-language captioning from MediaPipe Holistic keypoints, decoded by an LSTM action-detection model.

Read on IEEE Xplore →