Blog

Blog

Original writing on architecture, event-driven systems, reliability and engineering careers — drawn from production work on banking and high-traffic platforms, with the numbers and the trade-offs left in.

Clear

3 articles

Resilience4j, Kubernetes and Prometheus/Thanos: Engineering Toward 100% Resiliency

Circuit breaker states, autoscaling on breaker signals instead of raw CPU, Thanos for metrics that survive a Prometheus restart, and why Azure Monitor sees failures Resilience4j never will — from the work that took uptime from 99.5% to 99.95%.

Sep 21, 2026 By Vishnu Balachandran ResilienceCloud & KubernetesObservabilitySystem Design
👍 1

What Actually Breaks at 5 Million Requests a Day

Notes from building and hardening an API gateway handling 5M+ requests a day at 99.95% availability: the failure modes that do not show up until real traffic, and the patterns that actually held.

Sep 18, 2026 By Vishnu Balachandran API GatewayDistributed SystemsResiliencePerformanceSystem Design
👍 1

Progressive Delivery: Canary Releases, Feature Flags and A/B Testing

A deploy is not an event, it is a dial you turn. Building a canary framework with feature flags and A/B testing that cut deployment incidents 70% and doubled release velocity — and why the promotion gate matters more than the rollout mechanism.

Sep 16, 2026 By Vishnu Balachandran CI/CDCloud & KubernetesObservabilityResilience
👍 0