Insights on Agentic AI
Research, case studies, and field notes from the teams building production Agentic AI across Indonesia's banks, insurers, manufacturers, and public-sector institutions.
- Editor's Pick
Trade Finance Agentic AI: A Breakthrough 4 Days→15 Minutes Win with Redpumpkin AI
See how we fixed trade finance document processing Indonesia bottleneck, reducing processing time from 4 days to 15 minutes with agentic RAG.
April 23, 2026
All Articles
Filter by topic to see what's relevant to your team.
Model Distillation for Sales Analytics Workflows in Production
Model distillation is the right optimization path when sales analytics work has become repeatable. The CTO's task is to build a smaller student model around a narrow contract, then govern it like a production system.
August 4, 2026
LLM Quantization: GPU Memory Strategy for Enterprise AI Teams
Quantization is the first infrastructure move to test when GPU memory or serving cost blocks enterprise AI rollout. The production decision is about validating weight memory, activation math, KV cache growth, hardware kernels, and task-level quality together.
August 4, 2026
Speculative Decoding, Quantization, and Distillation Tradeoffs
Speculative decoding reduces generation latency when the workload is memory-bound. Quantization reduces model memory and serving cost by lowering numerical precision. Distillation creates a smaller model trained to imitate a larger one for a narrower workload. The right choice depends on whether the enterprise AI team is optimizing speed, cost, quality stability, or control.
July 24, 2026
Why Production AI Systems Need Test Harnesses Before Release
AI agents fail in production when only model outputs being tested. A production test harness evaluates the entire production environment together. This creates evidence for release decisions, regression control, and continuous improvement.
July 23, 2026

