Incident Manager Lead / AWS Cloud Operations
Peraton — Crystal City, VA
Dec 2019 – Present
- Lead major incident response across enterprise and AWS-aligned environments, coordinating infrastructure, security, application, and operations teams to minimize mission impact and restore service.
- Manage the incident lifecycle from triage and impact assessment through escalation, stakeholder communications, service restoration, root cause analysis, and post-incident review.
- Support mission-critical cloud operations across AWS Commercial and GovCloud environments with emphasis on availability, operational resilience, and federal security requirements.
- Develop and maintain incident-response playbooks and operating procedures integrating AWS CloudWatch, CloudTrail, Config, Systems Manager, ServiceNow, and BMC Remedy.
- Drive root-cause and recurring-incident analysis to identify operational trends, reduce repeat issues, and improve Mean Time to Resolution.
- Support monitoring, alerting, and escalation workflows using CloudWatch alarms, Lambda automation, EventBridge, and enterprise alerting capabilities.
- Collaborate with SRE, DevOps, security, network, and systems teams to manage SLAs, SLOs, service risk, and operational reliability.
- Support Infrastructure-as-Code practices using CloudFormation and Terraform to improve deployment consistency, rollback readiness, and repeatability.
- Provide executive and government stakeholders with incident reporting, operational trends, risk summaries, and service-availability updates.
- Mentor technical staff on incident-management practices, fault isolation, escalation procedures, and operational readiness.
- Participate in business-continuity and disaster-recovery planning and exercises to validate recovery procedures and mission readiness.