Full-Stack AI Engineer • India (Remote) • Open to Work
🔗 Portfolio • 📧 Email • 💼 LinkedIn • 🏆 LeetCode (Top 2%)
I am a Full-stack AI Engineer building production ML/LLM systems for finance & accounting. I combine modern AI application development (FastAPI, Next.js, LLMs, RAG, agentic workflows) with deep systems engineering (C++20, lock-free concurrency, sub-microsecond latency).
This dual-threat background enables me to build AI systems that are not just smart, but fast, reliable, and production-grade.
Important
Building production AI systems requires more than model prompting — it demands latency optimization, memory management, concurrency correctness, and infrastructure awareness.
My C++/quant background gives me this foundation. The same engineering discipline that achieves 270ns order-book latency (cache-line alignment, zero allocations, correct memory ordering) translates directly to optimizing LLM inference pipelines, building custom model servers, and debugging production AI latency issues.
(Note: Replace the placeholder [Add GIF here] text with actual image URLs of your apps once you record some 5-second GIFs!)
- Ultra-Low Latency LOB: Lock-free SPSC, slab allocator, 270ns latency, 2.5M msg/sec.
- Derivatives Pricing Engine: C++20, Pybind11, AVX2, 39.5x speedup vs QuantLib.
- LearnForge AI Course Player: Split-pane AI tutor with MDX lessons + live streaming chat.
(Example of how I design scalable AI solutions)
graph TD
A[Invoice Ingestion] -->|Raw PDF/Image| B(Tesseract OCR)
B -->|Extracted Text| C{3-Way Matching Engine}
C -->|Match| D[Approved Queue]
C -->|Mismatch / Anomaly| E(Ensemble Anomaly Detection)
E -->|Features| F((Isolation Forest))
E -->|Features| G((Z-Score))
F --> H[LLM Explainability Service]
G --> H
H -->|Human-readable alerts| I[Human-in-the-Loop Review UI]
- LLM Applications: NVIDIA NIM, Groq, OpenAI APIs, LangChain, SSE streaming, RAG, agentic workflows
- ML Models: Isolation Forest, Z-Score ensembles, XGBoost, BQML, scikit-learn
- Computer Vision: Tesseract OCR, multimodal vision LLMs (Llama-4-Maverick, Qwen2.5-VL)
- Backend: FastAPI (async, WebSocket, background tasks), Pydantic v2, SQLAlchemy, Alembic
- Frontend: Next.js 16 (App Router), React 19, TypeScript strict, Tailwind CSS, TanStack Query
- Databases & DevOps: PostgreSQL (Supabase), Redis, Docker, GitHub Actions, Vercel, Render
- Modern C++: C++20, Templates, concepts, atomics (
acquire/release), lock-free structures - Performance: Custom allocators (slab, arena), cache-line optimization, SIMD (AVX2),
perf
- jemalloc Contributor: Merged PRs #2815 (memory leak fix) + #2892 (lock-rank fix) in the memory allocator used by Meta, Firefox, and Redis.
- M.Sc Mathematics & Computing: Stochastic Calculus, Numerical Linear Algebra, Convex Optimization. CGPA 8.89/10 (Top 3%).
"The best AI engineers are systems engineers who understand that models are just one component in a production pipeline."
🎯 Current Status: Looking for a Junior Full-Stack AI Engineer role building production AI/ML features end-to-end.
🌍 Availability: Open to remote global opportunities (Visa sponsorship open, Indian Passport). 30 days notice.
💬 Let's Talk: nitinsaviobada@gmail.com


