Chaos Testing

Run chaos experiments without the guesswork

Most teams only discover resilience gaps during incidents. Harness injects controlled faults, measures what breaks with resilience scores and probes, and shows you exactly how to fix it — before your customers find out.

Systematically reduce MTTR

Make every incident your last

Replicate and prevent Harness runs a Chaos Experiment in staging that replicates the exact failure mode from your last incident, then adds it as a permanent regression test on every future deploy.

From firefighting to fireproofing Converts every production failure into a permanent test case so the same failure can never reach production twice. Engineering reclaims capacity for innovation.

Chaos Engineering

Stay one step ahead of every failure

Harness injects controlled faults across every layer of your stack — and tells you exactly what breaks before your customers find out.

Inject faults with a single click. Pod kills, network partitions, latency, and CPU/memory stress across Kubernetes, AWS, Azure, GCP, VMware, and bare metal.

ChaosGuard enforces safety gates. Blast radius limits and approval gates prevent experiments from causing unintended damage before any run begins.

Resilience probes run automatically. Validate system behavior throughout every experiment — no manual checks required.

230+ fault types. Comprehensive fault coverage across every layer of your stack and every major cloud and infrastructure platform.

Customer Success

Built to break. Proven to recover.

4 hrs → 2 min per test
We went from running fewer than one resilience test per month per application to 40 tests per month across 100 applications. What used to take 4 hours now takes 2 minutes. Harness gave our developers the self-service capability we needed to scale resilience testing across our entire migration.

, Director of Platform Engineering, Major US Airline

50% reduction in outages
Harness enabled us to automate disaster recovery testing and embed chaos experiments directly into our CI/CD pipelines. Our quality team went from 3 weeks of manual testing to automated, templated testing — and developers can now onboard and run approved tests themselves.

, Head of Site Reliability Engineering, Global Financial Institution

400% developer cost reduction
After migrating from Litmus to Harness Chaos Engineering, we cut cycle time in half and reduced the team required for resilience testing by 80%. Average test time dropped from 2 hours to 10 minutes — a 10x improvement that freed up engineers to focus on building, not testing.

, Principal Engineer, Platform Reliability, International Banking Group

Fault Library

230+ fault types for every failure scenario

From resource exhaustion to cloud outages to application-level failures — out-of-the-box experiments for every layer of your stack.

memory hog

Pod Memory HogNode Memory HogEC2 Memory HogECS Container Memory HogECS Fargate Memory HogLambda Update Function MemoryWindows EC2 Memory HogAzure Instance Memory HogVMware Memory HogVMware Windows Memory Hog

stop

ALB AZ DownCLB AZ DownNLB AZ DownAWS SSM Chaos By IDAWS SSM Chaos By TagEC2 Stop By IDEC2 Stop By TagECS Agent StopECS Instance StopECS Invalid Container ImageECS Task StopECS Update Container TimeoutLambda Delete Event Source MappingLambda Delete Function ConcurrencyLambda Function Layer DetachLambda Toggle Event Mapping StateLambda Update Function TimeoutRDS Instance DeleteRDS Instance RebootRDS Instance StopGCP VM Instance StopGCP VM Instance Stop By LabelAzure Instance StopAzure Web App StopVMware VM PoweroffVMware Host RebootVMware VM Service StopVMware Windows Service StopCF App Stop

I/O stress

Node IO StressEC2 IO StressECS Container IO StressAzure Instance IO StressVMware IO StressVMware Windows Disk StressDisk Fill

restrict access

Azure Web App Access RestrictECS Container Network RestrictionECS Network RestrictECS Update Task RoleLambda Update Role PermissionResource Access Restrict

disk loss

EBS Loss By IDEBS Loss By TagECS Container Volume DetachAzure Disk LossGCP VM Disk LossGCP VM Disk Loss By LabelVMware Disk Loss

time

Linux Time ChaosVMware Windows Time Chaos

network

EC2 DNS ChaosEC2 HTTP LatencyEC2 HTTP Modify BodyEC2 HTTP Reset PeerEC2 HTTP Status CodeWindows EC2 Blackhole ChaosWindows EC2 Network LatencyWindows EC2 Network LossECS Container HTTP LatencyECS Container HTTP Modify BodyECS Container HTTP Reset PeerECS Container HTTP Status CodeECS Container Modify HeaderECS Container Network LatencyECS Container Network LossPod Network LatencyPod Network LossPod Network CorruptionPod DNS ErrorPod DNS SpoofPod HTTP LatencyPod HTTP Status CodeEC2 Network LatencyEC2 Network LossAzure Instance Network LatencyAzure Instance Network LossLinux Network CorruptionLinux Network DuplicationLinux Network LatencyLinux Network LossLinux Network Rate LimitLinux DNS ErrorLinux DNS SpoofVMware Network CorruptionVMware Network LatencyVMware Network LossVMware DNS ChaosVMware HTTP LatencyVMware HTTP Reset PeerVMware HTTP Response ModifyVMware Network Rate LimitVMware Windows Blackhole ChaosVMware Windows Network CorruptionVMware Windows Network DuplicationVMware Windows Network LatencyVMware Windows Network LossNode Network LatencyNode Network Loss

kill

Container KillEC2 Process KillECS Container KillKubelet Service KillDocker Service KillLinux Process KillGCP VM Service KillVMware VM Process KillVMware Windows Process Kill

CPU hog

Pod CPU HogNode CPU HogEC2 CPU HogECS Container CPU HogECS Fargate CPU HogWindows EC2 CPU HogAzure Instance CPU HogVMware CPU HogVMware Windows CPU Hog

pod

Pod DeletePod AutoscalerTime ChaosPod API BlockPod API LatencyPod API Modify BodyPod API Modify HeaderPod API Status CodePod HTTP Modify BodyPod HTTP Modify HeaderPod HTTP Reset PeerPod IO Attribute OverridePod IO ErrorPod IO LatencyPod IO StressPod Network DuplicationPod Network Rate LimitPod Network PartitionPod Memory Hog Exec

kubelet

Node DrainNode RestartNode Taint

density

Pod CPU Hog ExecECS Task ScaleECS Update Container Resource LimitKubelet Density

linux

Linux CPU StressLinux Disk FillLinux Disk I/O StressLinux Memory StressLinux Service Restart

158 faults across 14 fault types

Get started with Harness Chaos Testing

See how Harness injects real faults into your stack, measures what breaks, and turns every incident into a test you'll never fail twice.