Third-year CS undergrad working on the unglamorous half of AI β the part where a model has to answer a real user, in real time, without falling over.
- ποΈ Building β a real-time voice RAG assistant on LiveKit + Sarvam, targeting sub-second turn latency
- β‘ Learning β TensorRT-LLM, Triton Inference Server, quantization trade-offs at serving time
- π Curious about β agent memory, retrieval that doesn't hallucinate, and where inference cost actually goes
- π¬ Ask me about β streaming STT/TTS pipelines, FAISS retrieval, FastAPI + WebSockets
- π« Open to β AI engineering internships, backend/SWE roles, open-source collaboration
|
Real-time conversational agent over a private knowledge base. Streaming audio in, grounded answers out. Pipeline β Silero VAD β Sarvam STT β FAISS retrieval β context injection β LLM β Sarvam TTS, streamed over LiveKit WebRTC.
|
Benchmarking what actually moves the needle on serving latency and throughput. Exploring β TensorRT-LLM engine builds, Triton model repositories, INT8/FP8 quantization, batching strategies, and the accuracy cost of each.
|
π‘ Swap these for your two strongest repos and add a one-line result to each β "cut p95 latency from 2.4s to 780ms" beats any badge on this page.
| Languages | |
| AI / ML |
|
| Backend | |
| Infra |
Training the model is the demo. Serving it at 3 AM to a user who doesn't care that it's AI β that's the engineering.
Building something in voice AI or inference? I'd like to hear about it.
ayushchougula.in Β Β·Β LinkedIn
