I build machine learning systems that connect models, data, APIs, and production infrastructure.
I am a Machine Learning Engineer and Data Scientist with experience across production ML platforms, predictive modeling, data science, distributed systems, and Generative AI.
My work spans both sides of machine learning: understanding data and building models, then designing the systems required to run AI reliably in production. I have worked with large-scale ML infrastructure handling 200M+ requests per day, operational and transactional datasets across the energy domain, and end-to-end AI applications using RAG, LangGraph, embeddings, vector search, FastAPI, and multimodal models.
I am particularly interested in GenAI, agentic AI, retrieval systems, recommendation systems, LLM evaluation, MLOps, and scalable AI infrastructure.
At Meta Reality Labs, I worked on large-scale infrastructure supporting ML artifact creation and AR/VR platform workflows processing more than 200M requests per day. My work focused on distributed orchestration, request identity normalization, Fetch or Create caching, idempotency, asynchronous validation, fault isolation, and production observability. I also supported the modernization of legacy build workflows through centralized resolution, feature flags, staged migrations, and safe rollout mechanisms across hundreds of build configurations.
At TCS, I worked with sales, order, logistics, and operational datasets across the energy domain. I used Python, SQL, statistical analysis, feature engineering, forecasting, and predictive modeling to study demand behavior, identify operational bottlenecks, detect anomalies, and support inventory and resource-planning decisions. I also worked on data-quality pipelines, near-real-time analytics, and Azure and Synapse based data workflows that helped move analytical outputs into scalable business and operational systems.
Mumbai, India | Jun 2018 to Mar 2020
I started my career working with order, dispatch, terminal, sales, and aviation fueling data across 15 fuel terminals. I analyzed more than 50,000 transaction records, performed exploratory and statistical analysis, engineered features, and investigated demand patterns, anomalies, and operational constraints. This work helped translate real operational problems into data-driven decisions and contributed to improvements associated with a 17% increase in fuel and lubricant sales.
An end-to-end AI system that transforms wardrobe images into structured garment intelligence and generates personalized outfit recommendations.
What I worked on:
- Multimodal garment understanding using image segmentation and CLIP embeddings
- Attribute extraction and structured garment metadata
- Embedding-based retrieval using pgvector and Supabase
- Personalized outfit ranking and recommendation logic
- RAG-based recommendation workflows
- Agentic orchestration using LangGraph
- FastAPI services for garment ingestion, retrieval, ranking, and recommendation delivery
- Feedback loops, fallback logic, duplicate detection, and recommendation evaluation
- LLM evaluation for relevance, groundedness, consistency, and hallucination risk
Tech: Python FastAPI LangGraph RAG CLIP PostgreSQL Supabase pgvector
Built an AI research agent capable of decomposing complex questions, retrieving relevant information, and generating responses while operating under strict context and token constraints.
The system combines agentic orchestration, vector retrieval, RAG, and summarization-based memory management to preserve useful context while keeping inference efficient.
Tech: Python LangGraph RAG Vector Search LLMs
Built a machine learning ranking system to predict user engagement and rank candidate items using contextual and behavioral signals.
The project covers the full ML lifecycle including feature engineering, model training, real-time inference, low-latency feature retrieval, event collection, retraining, and production-style serving.
Tech: Python LightGBM FastAPI Redis Kafka Airflow
Led an academic machine learning project focused on predicting high-risk clinical outcomes from structured healthcare data.
Evaluated multiple supervised learning approaches and improved key classification metrics including precision, recall, and ROC-AUC by approximately 30%.
Tech: Python Scikit-learn Pandas Machine Learning
Created a classification pipeline using customer engagement and behavioral features to identify customers at risk of churn, achieving approximately 92% classification accuracy.
Tech: Python Scikit-learn Pandas
LangGraph LangChain RAG LLM Inference OpenAI API Hugging Face Transformers VLM Integration
PyTorch TensorFlow Keras XGBoost LightGBM Scikit-learn Feature Engineering Model Evaluation
CLIP Embeddings Vector Search pgvector Supabase Embedding-Based Retrieval
Python FastAPI Docker REST APIs Real-Time Inference Model Deployment Observability
SQL PostgreSQL Azure Databricks Synapse Analytics Blob Storage Logic Apps
Git API Integration Distributed Systems Production ML Infrastructure Pandas
University of Houston, C. T. Bauer College of Business CGPA: 3.83 / 4.0
Google Cloud Professional Machine Learning Engineer
I enjoy working on problems involving:
- Generative AI and RAG
- AI Agents and Agentic Workflows
- Applied Machine Learning and Data Science
- Recommendation and Ranking Systems
- Embeddings and Retrieval
- LLM Evaluation and Reliability
- MLOps and ML Platforms
- Large-Scale AI and Data Infrastructure
- LinkedIn: linkedin.com/in/santhosh-botcha
- Email: santhoshbotcha97@gmail.com
- GitHub: Explore my repositories and projects here
- Phone: (346) 438-0655