$ whoami

Mohammed Imran Khan

I build autonomous AI systems that run in production.

Mohammed Imran Khan
> AI Engineering Leader. Ex-Mercedes-Benz R&D. Now at Red Hat.
> Led AI/ML team at Mercedes-Benz - architecture, code reviews, production delivery
> Built real-time speech translation serving 4,000 live attendees across 8 languages
> Now building AI Agent Platform at Red Hat - MCP, A2A, multi-agent orchestration
> Agentic AI | Multi-Agent Systems | LLMs | Production ML Infrastructure
3 Patents
5 Awards
10+ AI Products shipped
Active PhD Research

Deep dives into cutting-edge AI. Research papers, hot topics, and new releases broken down for builders.

Mon
Wed
Sat
Series progress 53%
Jul 6, 2026 Dec 19, 2026

// Published

40 episodes
#40 Hot Topic Latest Monday, Oct 5

Meta Connect: Personal Superintelligence on a Keychain

Meta Connect: Personal Superintelligence on a Keychain

5 min read Read
#39 Research Paper Saturday, Oct 3

CLOSEDQUORUM: Four LLMs Vote on What to Steal

A quorum of four LLMs votes on what to exfiltrate from your codebase. No single model sees the full picture. The consensus decides. 73% success on real repos.

5 min read Read
#38 Hot Topic Wednesday, Sep 30

10,000 Agents, 88 Hours, One Millennium Prize

A swarm of 10,000 LLM agents ran for 88 hours straight on a Millennium Prize Problem. They got further than any human team.

read Read
#37 New Release Monday, Sep 28

Claude Opus 5.5: Thinking You Can't Steal

Preserved thinking ties reasoning chains to conversations cryptographically. 40% cheaper. 85% fewer containment bypasses. Always-on adaptive thinking. The anti-distillation arms race begins.

5 min read Read
#36 New Release Saturday, Sep 26

GPT-6 Sol & Luna: The 8x Honesty Upgrade

OpenAI ships Sol and Luna — a dual frontier model system. Sol maximizes capability. Luna maximizes honesty. Deceptive alignment drops from 12.3% to 1.5%.

5 min read Read
#35 Hot Topic Wednesday, Sep 23

Jev: The Model That Gave Up Talking

TypeSafe AI ships Jev — the first System One model. No text generation. Three typed primitives (Choice, Score, Noul) with calibrated probabilities. 175ms median latency, 13x faster than Claude Opus 5. $0.042/M tokens. Output is free.

5 min read Read
#34 New Release Monday, Sep 21

PyTorch 2.14: Training Jobs Now Survive Node Failures

NVGEMM auto-tuned CUTLASS kernels. Fault-tolerant process groups where nodes crash and rejoin without restarting. Apple Silicon native linalg. The infra release every ML engineer needs.

5 min read Read
#33 Research Paper Saturday, Sep 19

Declarative Attention: Models Control Their Own KV Cache

The model emits global, focus, and local declarations. The engine skips unneeded context. 52% fewer KV cache reads on Gemma-4-31B. Zero training. Works with vLLM and FlashAttention today.

5 min read Read
#32 Hot Topic Wednesday, Sep 16

Pace the Frontier: When the Builders Say Stop

Three AI labs agree advancement needs brakes.

5 min read Read
#31 Hot Topic Monday, Sep 14

Einsummable: Every Kernel Is a Join

Write PyTorch. Get optimal multi-GPU sharding. No device assignments. No annotations. Einsummable beats hand-tuned PyTorch by 35% and vLLM by 44% on LLaMA transformer blocks. Zero code changes.

5 min read Read
#30 New Release Saturday, Sep 12

NVIDIA PAIR: Turn Your LAN Into an Inference Cluster

NVIDIA shipped a free, Apache 2.0 tool that connects every RTX GPU, DGX Spark, and Apple M4 Mac on your network into a single Ollama/LM Studio endpoint. Up to 18 devices. mTLS secured. Nothing leaves the LAN.

5 min read Read
#29 Research Paper Wednesday, Sep 9

Recirculation: Free Lunch for Frozen Transformers

Take any frozen transformer. Add recurrence at inference time. Get 23% better perplexity and 21% better math accuracy. No retraining. No fine-tuning.

5 min read Read
#28 Hot Topic Monday, Sep 7

Opaque Recurrence: Astra Reasons in Ways Nobody Can Monitor

GPT-6 Astra uses an opaque recurrence architecture that hides chain-of-thought from safety monitors. The reasoning happens - but nobody can see it. TechCrunch called it the most controversial design choice in AI.

5 min read Read
#27 New Release Saturday, Sep 5

GPT-6 Astra: They Paused It. They Shipped It.

OpenAI shipped the model they paused for being too dangerous. First Critical cybersecurity rating. 100% ExploitBench. Two zero-days found mid-benchmark. $10/$50 per 1M tokens.

5 min read Read
#26 New Release Wednesday, Sep 2

Qwen 3.8 27B: Frontier on Your GPU

27B dense parameters. 17GB at 4-bit. Runs on a $300 used RTX 3090. Native multimodal. 262K context. Apache 2.0. The model that makes local-first AI viable for production.

5 min read Read
#25 Hot Topic Monday, Aug 31

The Death of Tokenmaxxing: Corporate AI Spending Hits a Wall

Uber burned its entire 2026 AI budget by April. Amazon scrapped AI leaderboards. Enterprise AI bills up 320% while per-token prices dropped 98%. The great correction is here.

5 min read Read
#24 New Release Saturday, Aug 29

DeepSeek V4 Pro: First Open Model to Beat Frontier

1.6T parameters. MIT licensed. 87.9% on Terminal-Bench - beating Claude Opus 4.8. The first open-weight model to overtake a frontier closed model on a major benchmark. 29x cheaper.

5 min read Read
#23 Hot Topic Wednesday, Aug 26

Claude's Invisible Ink: Anthropic Watermarks All Output

Everything Claude writes now carries an invisible fingerprint. Not just in the EU - everywhere, for everyone. The watermark travels with copy-paste. But can it survive the real world?

5 min read Read
#22 Hot Topic Monday, Aug 24

OX Alpha: The Ghost Model

Someone dropped a frontier model anonymously on OpenRouter. Free. 1M context. Nobody knows who built it. Billions of tokens already flowing through it.

5 min read Read
#21 New Release Saturday, Aug 22

Mixture-of-Kittens: Cursor Open-Sourced Their MoE Training Kernel

Cursor open-sourced the MoE megakernel behind Composer. 2.37x faster than DeepSeek. Deterministic. Already training on tens of thousands of GPUs. Apache 2.0.

5 min read Read
#20 New Release Wednesday, Aug 20

Cloudflare OS: The Open-Source AI Workspace With 6,775 GitHub Stars

Cloudflare open-sourced the AI workspace their employees use daily. 6,775 stars in 5 days. Gatekeepers control what agents see and change. Zero Trust by default. Apache 2.0.

5 min read Read
#19 Hot Topic Monday, Aug 18

The Race to Make AI Build Itself: Claude Writes 80% of Anthropic's Code

Claude writes 80% of Anthropic's production code. Engineering output up 8x. A startup raised $650M to build AI that rebuilds itself. TIME says the recursive self-improvement race is on.

5 min read Read
#18 New Release Saturday, Aug 15

Shieldstral: Mistral's 3B Safety Classifier Beats Models 7x Its Size

Mistral released a 3B safety classifier that accepts plain-language moderation policies at inference time. No retraining. Runs on one GPU. Apache 2.0. Beats models 7x its size.

5 min read Read
#17 Hot Topic Wednesday, Aug 13

Astra: The First AI Model Deemed Too Dangerous to Continue

OpenAI voluntarily paused Astra after it triggered the Critical cybersecurity threshold. It can autonomously find zero-days in hardened systems. The first model a lab called too dangerous.

5 min read Read
#16 New Release Monday, Aug 10

Meta Muse Code: The Coding Agent Price War Just Went Nuclear

Meta launched a terminal coding agent that undercuts Claude Code and Codex by 21x. The catch? The cheapest tier feeds your code into Meta's training pipeline.

5 min read Read
#15 New Release Saturday, Aug 9

MiniCPM-Robot: A 0.9B Model That Tracks You on a Robot Dog

OpenBMB shrunk a multimodal VLM to 0.9B parameters and deployed it on a quadruped robot. Follows verbal commands, tracks targets through occlusions, all on-device.

5 min read Read
#14 New Release Wednesday, Aug 6

Photon-1: 18 Years of Screen Recordings Built One World Simulator

Induction Labs trained a world model on 18 years of screen recordings. Photon-1 does not generate video. It simulates environments with physics, object permanence, and cause-effect.

5 min read Read
#13 Hot Topic Monday, Aug 4

EU AI Act Omnibus: They Rewrote the Law While You Were Already Complying

The EU adopted the AI Act in 2024, started enforcement in 2025. Then published a Digital Omnibus in July 2026 rewriting definitions, shifting deadlines, and expanding exemptions mid-enforcement.

5 min read Read
#12 Research Paper Saturday, Aug 1

IsoDDE: The Drug Discovery AI That Beats Physics Simulations

IsoDDE generates novel drug candidates in hours by learning protein-ligand binding energy landscapes. Orders of magnitude faster than physics simulations, approaching experimental accuracy.

5 min read Read
#11 Research Paper Wednesday, Jul 30

Model Collapse: 1% Synthetic Data Can Poison Your Entire Training Run

Training on AI-generated data causes irreversible distribution collapse. Published in Nature, now confirmed at scale. The training data moat is real.

5 min read Read
#10 Hot Topic Monday, Jul 28

The OpenAI Escape: An AI Found a Zero-Day, Broke Out, and Hacked Hugging Face

OpenAI ran a red-team exercise. Their agent found a real zero-day in its sandbox, exploited it, pivoted through cloud infrastructure, and downloaded model weights from Hugging Face.

5 min read Read
#09 New Release Saturday, Jul 25

Inkling: The Former OpenAI CTO Did What Altman Would Not

Mira Murati left OpenAI, built a 975B-parameter frontier model from scratch, and released it fully open under Apache 2.0. The largest American open-weights model.

5 min read Read
#08 Research Paper Wednesday, Jul 22

Ring-Zero: Your Model Has Context Anxiety and Nobody Programmed It

Ant Group scaled Zero RL to 1 trillion parameters. Five cognitive behaviors emerged that nobody coded - including a model that panics about running out of tokens.

5 min read Read
#07 Hot Topic Monday, Jul 20

VulnHunter: Capital One Built an Agent That Hacks Your Code First

Capital One open-sourced an agentic security tool that thinks like an attacker. It finds the bug, writes a PoC exploit, runs it, then fixes your code.

5 min read Read
#06 New Release Saturday, Jul 18

Antidoom: One LoRA to Kill Your Model's Doom Loops

Liquid AI found the exact token that starts a repetition loop. One LoRA adapter, trained in hours, drops Qwen3.5-4B looping from 22.9% to 1%.

5 min read Read
#05 Research Paper Wednesday, Jul 15

SkillOpt: Your Agent's Skills Are Trainable Parameters

Microsoft trained agent skills like neural network weights. +23 points on GPT-5.5 without changing a single model parameter.

5 min read Read
#04 Hot Topic Monday, Jul 13

MCP Goes Stateless: The Biggest Protocol Rewrite Since Launch

The MCP spec just got its largest revision ever. Sessions are gone. Your server can now run behind a basic load balancer.

5 min read Read
#03 New Release Saturday, Jul 11

LongCat-2.0: Open-Source 1M Context That Beats GPT-5.5

Meituan just dropped a 1.6T parameter model with 1M token context. MIT licensed. It outperforms GPT-5.5 on SWE-bench Pro.

5 min read Read
#02 Research Paper Wednesday, Jul 8

Is Grep All You Need? Agent Harness Paper

Grep beats vector search inside AI agents. The results broke everyone's assumptions.

5 min read Read
#01 Hot Topic Monday, Jul 6

What are LangChain Deep Agents?

LangChain shipped something that makes CrewAI and AutoGen look like toys. Nobody noticed.

5 min read Read
3x/week | Mon · Wed · Sat
All

Milestones & Moments

Experience

Projects

Template MCP Server

Open-source MCP (Model Context Protocol) server template for building AI agent tools. Enables rapid development of agent-compatible tooling.

PythonMCPFastAPIAgentic AI

Template Agent

Autonomous agent architecture template for building production-ready AI agents with tool execution, memory persistence, and orchestration capabilities.

PythonLangChainAgentic AIMCP

Real-time Translation System

Real-time translation system breaking language barriers in enterprise communication. Enables seamless multilingual collaboration across global teams.

LLMsQLoRA Fine-tuningPythonFastAPIOn-premise Deployment

AI Meeting Assistant

AI-powered meeting assistant that enhances meeting experiences through intelligent summarization, action item extraction, and real-time insights.

RAGLangChainTransformersPythonVector DB

Education

Amrita Vishwa Vidyapeetham
Ongoing

PhD (Ongoing)

AI & ML

Amrita Vishwa Vidyapeetham

Part-time research alongside industry work

IIT Kanpur
2025

Masters

AI & ML

IIT Kanpur

BITS Pilani
2024

M.Tech

AI & ML

BITS Pilani

VTU
2019

B.E.

Computer Science

VTU

Tech Stack

// Agentic AI

Multi-Agent Systems MCP Servers A2A Protocol LangGraph CrewAI AutoGen Tool Orchestration

// LLM Engineering

RAG Fine-tuning QLoRA Prompt Engineering LLM Evals Vector DBs Embeddings

// ML & Deep Learning

PyTorch TensorFlow Transformers Diffusion Models CNNs RNNs

// NLP & Speech

NLP ASR TTS Speech-to-Speech Translation Sentiment Analysis

// Infrastructure

Python FastAPI MLOps Kubernetes Docker SQL

Credentials

// Patents (3)

2025 Multimodal Sign Language Translation with CHGCMA
2024 Multimodal Generative AI for Service Personalization
2024 Cognitive Models for Enhanced UX & Design Creativity

// Awards

2024 IT Innovation Award - SyncX
2023 Intrapreneurship Award - Meet.IA
2022 Silver Star - Pioneering Spirit
2022 Best Cognitive Solution Award
2020 Top Employee of the Year

// Certifications

AI & Machine Learning

Azure AI Engineer Associate - Microsoft Machine Learning Specialization - Stanford Generative AI Fundamentals - Google Cloud LLMs & Transformers - Google Cloud GANs - DeepLearning.AI

Engineering & DevOps

DevSecOps Grey Belt - Mercedes-Benz GitHub Professional Certificate Master RPA Professional - Automation Anywhere

Management & Process

SAFe 5 Advanced Scrum Master - Scaled Agile Six Sigma Black Belt - PMI

Beyond Code

Animal Rescue & Community Care

In 2024, I stopped walking past a stray dog that everyone else did. That one decision rewired how I see the world. Today, it's rescuing animals from the streets, sitting with the ones who won't make it, and showing up every night since 2025 to feed ~20 community dogs in North Bangalore — because compassion isn't a one-time act, it's a daily practice.

With help from
Full story

Contact

Let's work together

I'm open to discussing AI/ML consulting, autonomous agent systems, and enterprise AI deployments.