Technical articles from production AI work and academic study.
A $7K datacenter L40S lost to a $1.6K RTX 4090 on Whisper inference. The GPU was innocent: INT8 quantization was the culprit. How I traced it, suspect by suspect, and the bandwidth math behind it.
Read on Medium →Algorithms are averaging machines: they know least about career changers, immigrants, and anyone off the average. An essay from my CAS ethics module on why “fair” algorithms still miss the people who need fairness most.
Read on Medium →H100s, NIM microservices, and Milvus on Kubernetes, until an unauthorised registry key left every pod stuck in ImagePullBackOff. How I patched the stack back to life.
Read on Medium →Passed on the first attempt with about ten hours of prep. What worked: trusting hands-on experience, reading the official docs, and practising in labs.
Read on Medium →Empty streams from an LLM inference endpoint, and for once the bug wasn’t in my code. How I ruled out my service, reported the issue to Postman, and saw it fixed two weeks later.
Read on Medium →CAS final project (University of Bern, 2026): open-source LLMs on BIRD text-to-SQL with agentic self-correction.
Final revision in progress. Subscribe to get the paper when it's published.
Field notes from production AI: what breaks, what holds, and what it teaches. One email when something ships.
Double opt-in. Unsubscribe anytime. Privacy notice · Powered by Buttondown