Podcast
AI Post Transformers
AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis,…
Already have it? Open in the app
Latest episodes
- SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators11 Aug 2026
- Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing10 Aug 2026
- MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA10 Aug 2026
- Silicon Showdown: GPU vs Apple Silicon LLM Inference Limits9 Aug 2026
- Making Every Verified Token Count in MoE Speculative Decoding4 Aug 2026
- Model Predictive Control's Real-Time Structure, from Chapter to Cockpit4 Aug 2026
- Global Memory Bloat in Long-Context LLM Serving4 Aug 2026
- FreeAct: Rethinking One-to-One Transforms for LLM Quantization4 Aug 2026
- DualDecoder: Fixing GPU Memory Bloat in Long-Context KV Cache Offloading4 Aug 2026
- Adapting Without Forgetting: A Lifelong Learning Roadmap for LLM Agents2 Aug 2026