Senior AI Engineer — Agentic Systems, Voice AI, GPU

Chandan Kumar

7+ years building production software. Now shipping low-latency voice agents, agentic developer tooling, and LLM fine-tuning.

5×
TTS speedup
Qwen3-TTS megakernel
47ms
Time-to-first-audio
RTX 5090
$10M+
ARR-worth initiatives
BrowserStack

About

Engineer-first AI builder

Senior software engineer with 7+ years across JPMorgan Chase and BrowserStack, now building production AI — agentic developer tooling, low-latency voice agents, and LLM fine-tuning. Comfortable from GPU-level CUDA work to product delivery.

B.Tech, Electrical Engineering from IIT (BHU) Varanasi (CGPA 9.03). Mentor of two engineers and an AI-SDLC SPOC for Growth Engineering.

BrowserStack AI Champion 2025

Featured Projects

What I'm shipping

Qwen3-TTS Megakernel

Persistent-kernel CUDA megakernel repurposed as the talker decoder for Qwen3-TTS (0.6B). 47 ms median TTFB and 0.145 RTF on RTX 5090 — 5× faster than the stock pipeline. Diagnosed and fixed a grid-barrier race deadlock; verified numerical parity against HF reference.

CUDA Qwen3-TTS bf16
View on GitHub →

Real-Time Voice Agent

Full WebSocket voice agent (mic → faster-whisper STT → gpt-4o-mini → megakernel TTS → speaker) on Pipecat with Silero VAD endpointing, barge-in, and 9 switchable voices. Cut speech-end-to-first-reply latency 2–3.7× to 0.7–1.1 s by colocating STT on the GPU.

Pipecat Whisper Silero VAD
View on GitHub →

SDD Context Engine

Python FastMCP server exposing custom tools (feature-knowledge queries, dev-phase activation, mocking helpers) that injects codebase-grounded context into Claude Code / Copilot. Drives the org-wide 6-phase agentic Spec-Driven Development pipeline at BrowserStack.

FastMCP Python Claude Code
BrowserStack — internal

Indic LLM Fine-Tuning

Fine-tuned open-source LLMs (Qwen) for Indic languages using QLoRA. Evaluated tokenizers and compared SFT vs. continued pretraining on Indic corpora.

QLoRA Qwen SFT
View on GitHub →

Experience

Where I've worked

Senior Software Engineer — BrowserStack

Oct 2024 — Present
  • Drove org-wide adoption of a 6-phase agentic Spec-Driven Development pipeline — cutting dev time ~50% and story bugs up to 70% on new initiatives.
  • Built the SDD Context Engine (Python FastMCP) and owned an agentic AI harness with 11 skills and 7 phase-specific agents.
  • Led AI-driven Nightwatch → Playwright test migration; shipped AI code review as a deploy gate.

Software Engineer — BrowserStack

Mar 2022 — Sep 2024
  • Built a self-serve promotional engine and upsell workflows across Pro plans; led $10M+ ARR-worth retention and expansion initiatives.
  • End-to-end marketing-attribution system (UTM tracking, cross-page activity, Salesforce sync).

Associate Software Engineer — JPMorgan Chase

Mar 2020 — Mar 2022
  • Built a Spring Boot + React monitoring and alerting tool tracking 26 applications across stacks; mentored an offshore team for ~1 year.

Software Engineer (SDE 1) — JPMorgan Chase

Jul 2019 — Mar 2020
  • Built an end-to-end Spring Boot application for the SPG mortgage-analytics platform — the entry and exit point for all securitized-product analysis at JPMC.

Writing

Notes from the build

Long-form posts on agentic SDLC, voice-AI engineering, GPU rabbit holes, and building with LLMs. Numbers, code, and what didn't work.

Loading latest posts…

Read the blog →

Content

I teach AI system design

I create AI system-design education for engineers on Instagram and YouTube — multi-agent patterns, LangChain / LangGraph / CrewAI, distributed systems, and database internals.

Skills

What I work with

AI / GenAI
LLM fine-tuning (QLoRA, SFT) RAG Multi-agent systems Agentic SDLC AI code review Prompt & context engineering Agent evaluation
Agent Frameworks
LangChain LangGraph CrewAI MCP / FastMCP
Voice AI
Pipecat Twilio ElevenLabs Cartesia Whisper / faster-whisper Silero VAD Qwen3-TTS
Languages
Python Java (Spring Boot) JavaScript / React C Bash SQL
Systems & Infra
CUDA / GPU programming FastAPI Docker PostgreSQL Redis Kafka MongoDB AWS

Contact

Let's talk

Open to senior AI / GenAI roles, voice-agent work, and speaking. Fastest reply via WhatsApp.