Designed and implemented an event-driven, automated self-healing pipeline for a microservice deployed on a Kubernetes cluster. The system eliminates manual intervention during application failures by automatically detecting HTTP 500 error spikes and triggering an automated remediation process to restore cluster health.
Impact: Built a fully automated Kubernetes incident response system using Prometheus and Ansible, reducing MTTR to < 1 minute.