I'm the founder and CEO of Vitalops, where I run everything on the ML engineering side. The work I keep for myself is inference: tuning engines until the tokens-per-second number actually moves, then proving with evals that the accuracy survived.

Most recently I served a consumer AI product with more than 10 million users: patching and tuning the vLLM and SGLang serving stack on NVIDIA Blackwell GPUs until it ran at over twice the benchmark speed, sustaining 10M+ tokens per minute under replayed production load and placing top 3 on Artificial Analysis, with custom eval suites proving the accuracy held. At Vitalops I hire and lead the engineering team, set the research direction, and still own the inference stack end to end: kernels, serving, quantization, evals.

Before founding, I spent years in AI research and ML engineering, with the papers on Google Scholar. I contribute to high-impact open source: merged patches in vLLM, roboflow/supervision, Lightning Flash and Talos, alongside the projects I maintain myself, all on github.com/abhijithneilabraham.

Resume (PDF)

Abhigith Neil Abraham

⚡ What I work on

  • Inference performance. Making a model serve more tokens per second on the hardware you already have: speculative decoding, quantization, kernel-level fixes, GPU profiling and load testing, with an eval harness alongside so the speedup is never paid for in accuracy.
  • Agent harnesses. Coding agents and multi-agent systems that hold up on real benchmarks: verify-then-escalate chains, LLM-as-judge verification, tool and MCP design, context engineering, failure-mode analysis.
  • Data at scale. LLM transformations over datasets far larger than any context window, the subject of Datatune and of my DATA 2026 paper.

📮 Writing

Sep 28, 2026Learning Inference: a case study for beginners
Sep 25, 2026Learning LLM Inference: confidential computing to keep data private
Sep 22, 2026Learning LLM Inference: scaling from a single node to millions of users
Sep 19, 2026Learning Inference: how to host and improve the token speed of an LLM
Aug 8, 2026A simple practical mental model for LLM inference optimization
May 27, 2026Solving your FOMO in this agentic AI world
Jun 5, 2025Solving complex data pipelines with Composio + Datatune
Apr 24, 2024A very basic roadmap for starting with machine learning and generative AI
Apr 4, 2024Solving your FOMO about everything in LLMs
Mar 27, 2024Data for LLMs: navigating the LLM data pipeline

Earlier pieces for Paperspace and E2E Networks, and the rest on Medium. Everything I write is indexed on the literature page.

📦 Open source

  • Added KeyPoints.merge(); fixed process_video hanging past the end of a video, out-of-bucket detections in size-bucketed metrics, crop/overlay annotators drawing outside the scene, and empty/numpy-index handling in key points.
  • The community hardware plugin for vLLM on Apple Silicon. Metal attention-kernel fixes (unclamped running max on prefill and window rows, neutral window-skipped partitions in the split-K reduce), TurboQuant K block scales, GGUF loader dtype casting, KV-budget validation, MLX cache synchronization in profile_run, and replacing VLLM_METAL_MEMORY_FRACTION with --gpu-memory-utilization.
  • Added the TextEmbedder task (sentence-transformer embeddings) and its docs.
  • autonomio/talos5 merged PRs
    Implemented DistributeScan / RemoteScan for distributed hyperparameter search across machines, later spun out as Jako.
  • Openvibe1.5k+ stars
    Agentic coding harness.
  • TableQA320+ stars
    Natural language to SQL over tabular data.
  • DatatuneVitalops
    Large-scale LLM-powered data transformation engine.
  • OpendeskVitalops
    Multi-machine computer-use MCP server.
  • EasyEnergyAutonomio
    Carbon and energy tracking for ML experiments.
  • JakoAutonomio
    Distributed hyperparameter optimization across machines.

📄 Research

Full list on Google Scholar.

🎤 Talks

More on YouTube.

🎓 Teaching

🧰 Toolkit

LLM inference & performance
Inference optimization, speculative decoding, quantization, GPU profiling & bottleneck analysis, confidential computing, load testing · vLLM, SGLang, MLX, NVIDIA Blackwell GPUs
Agent systems
Agent harness design, multi-agent orchestration, verify-then-escalate chains, LLM-as-judge verification, tool & MCP design, context engineering, failure-mode analysis
Evaluation & research
Eval harness design, benchmark replication, experiment design & ablation, statistical significance testing, accuracy-regression verification · Terminal-Bench, BFCL
Systems & delivery
Distributed GPU infrastructure, full-stack engineering, large-scale data transformation, secure deployment & attestation · Python, SQL, Postgres, Docker, AWS, GCP, Linux

BTech in Electrical and Electronics, College of Engineering Trivandrum, 2016–2020.

📬 Follow me

The best way to reach me is email. I'm also on X, GitHub and LinkedIn. Always open to interesting conversations and collaboration.