Nisarg
Kudgunti
AI Engineer who ships multi-agent systems for Fortune 500 clients by day, and built a transformer from raw matrix multiplications by night.
Currently architecting AI vulnerability remediation and RAG systems as an Associate AI Engineer at Techolution in Hyderabad. Before that, an AI Intern at the same company, and a Data Science Intern at AlgoAnalytics. Two years spent in production GenAI. One side project spent rebuilding Llama 3's architecture from scratch, just to know exactly what happens inside it.
· 01 / SUBJECT
About
Nisarg builds the infrastructure layer of generative AI: multi-agent orchestration, retrieval pipelines, and the evaluation loops that keep them honest in production. At Techolution, his work spans vulnerability remediation agents, hybrid RAG systems over thousands of enterprise forms, and natural-language-to-SQL pipelines that turned multi-day reporting into a sub-minute process.
Most of that work means directing large pretrained models. To understand what is actually happening underneath them, he implemented the Llama 3 architecture from first principles in PyTorch: rotary position embeddings, RMSNorm, SwiGLU, the full decoder stack, then pretrained and fine-tuned it himself. That same instinct toward verifiable AI runs through the rest: Counterpoint traces every claim it makes back to a verbatim source quote, and APEX puts a deterministic code-level verifier between its F1 assistant and the user, so an unciteable answer is hedged rather than shipped. All are on file below.
He holds a BE in Computer Engineering with Honors in AI/ML from PVGCOET, Pune, and is Microsoft Certified in Azure AI Fundamentals.
Hons. AI/ML, PVGCOET
PyTorch, FastAPI
· 02 / RECORD
Work on file
- Architected a multi-agent vulnerability remediation system for a Fortune 500 media client, integrating Checkmarx PR and push webhooks to detect and fix SAST, SCS, and IaC issues across 50 repositories, with threshold-gated parallelism per vulnerability type.
- Built a repo-scoped feedback loop on an AlloyDB vector store, embedding accepted, regenerated, and PR-comment feedback with language and type metadata, retrieving the top 5 fixes by cosine similarity to keep outputs aligned to each codebase.
- Built a hybrid RAG conversational agent over 4,000+ cost estimation forms for an enterprise retail client, using whole-form chunking and LLM-based entity extraction to pre-filter retrieval by known form fields.
- Designed a key-aware citation system that highlights the exact field and value a generated answer was sourced from, directly inside the original form.
- Reduced construction schedule reporting from 2 to 3 days down to under a minute with an NL-to-SQL agent over 8,000+ forms, consolidating 80+ tables into 20 UI-mirrored tables to cut hallucination, and a multi-tool pipeline of table selector, SQL writer, validator, executor, and LLM summarizer with table-level schema metadata, DML blocklist validation, cross-project comparisons, trend analysis, and XLSX export.
- Building a multi-agent platform reducing client project delivery from 3 to 4 weeks down to 3 to 4 days via automated data categorization and segmentation, with gap resolution and impact analysis that identifies low-confidence AI outputs, guides structured resolution, and applies fixes across dependent segments.
- Built Project Analyser, a 10+ language GitHub repository analysis pipeline benchmarked at 80k files against the Linux kernel, using Qwen on Cloud Run via vLLM for per-file summarization and Gemini on Vertex AI for executive summaries, service extraction, and dependency graphs.
- Extended the pipeline with HDBSCAN SOM-based clustering to derive hierarchical feature taxonomies and LLM-generated JSON schemas rendered as ReactFlow architecture diagrams.
- Reduced VertexAI inference cost by more than 90% by replacing proprietary LLMs with Qwen-2.5-Coder served on Cloud Run via vLLM.
- Enhanced a RAG application using LangChain, FAISS, LLM reranking, and innovative chunking strategies to improve RAGAS metrics and Generative AI performance.
· 03 / EXHIBITS
Case files
· 04 / INDEX
Working stack
· 05 / ADDITIONAL RECORDS
Also on file
· 06 / REQUEST ACCESS
Get in touch
Open to AI Engineer, ML Engineer, GenAI Engineer, and Applied Scientist roles, in Pune, Bengaluru, Hyderabad, or remote.