MinersMe Cloud logoMinersMe Cloud Create account →

← Blog · Guides & insights · August 4, 2026

Guide to Mining Incident Response for ASIC Fleets

Guide to Mining Incident Response for ASIC Fleets

A 3 MW hall can lose revenue long before anyone calls it an outage. Hashrate drifts downward, intake temperatures creep up, a breaker load approaches its limit, and workers begin rejecting shares. Then a technician notices the alert after a shift change. A guide to mining incident response has to start before the fleet is visibly down, because the fastest recovery is the one that contains a fault while it is still small.

Mining incidents are not generic IT tickets. A bad pool endpoint can redirect a customer’s hashpower. A failing fan can turn into board damage. A tripped breaker can take down hundreds of machines, while an uneven restart can trip it again. The job is to protect people and electrical infrastructure first, then protect production, evidence, customer commitments, and the next shift from repeating the same failure.

What Counts as a Mining Incident?

Treat an incident as any condition that threatens safe operation, expected hash production, customer revenue, or control of the fleet. That includes a site-wide utility event, but it also includes the quieter failures that compound across a large deployment: a rising count of missing hashboards, repeated thermal throttling, an abnormal reject rate, offline miners clustered on one PDU, unauthorized pool worker changes, or a maintenance queue that is no longer moving.

The right response depends on blast radius and rate of deterioration. One miner with a weak board is maintenance work unless it is exposing a power or fire risk. Fifty miners going offline on the same rack is an infrastructure investigation. A pool change across a customer group is a security incident until proven otherwise. Do not wait for a single universal severity definition. Build thresholds around the conditions that actually cost your operation money or create danger.

A useful severity model has four levels:

Severity must be allowed to move. Ten isolated fan faults can become a high-severity cooling or dust problem when they share a location, timing, or equipment batch.

Guide to Mining Incident Response: The First 15 Minutes

The first 15 minutes determine whether the team is operating from evidence or from noise. The incident lead should establish ownership immediately. One person coordinates decisions and communications; technicians execute assigned checks; everyone records timestamps and observations in the same incident record. A crowded chat room is not a command system.

1. Make the area safe and stop the spread

If breakers are tripping, temperatures are outside safe limits, or there is any sign of arcing, smoke, damaged cable, or abnormal electrical odor, isolate the affected equipment under site safety procedures. Do not keep cycling breakers to chase uptime. Repeated energization can turn a recoverable fault into damaged distribution gear or a personnel hazard.

For thermal events, determine whether the issue is local airflow, fan failure, a failed extraction path, ambient conditions, or a broader cooling-system problem. Power down only the affected miners when possible, but do not hesitate to take a rack or container offline if continuing operation risks equipment damage.

2. Define the blast radius from telemetry

Start at fleet level, then descend. Compare site hashrate against expected output, offline count by container and rack, breaker load by circuit, inlet and outlet temperatures, reject rate, pool connectivity, and worker configuration changes. The goal is not to collect every metric. It is to identify the shared boundary.

If offline miners align to one breaker, inspect the electrical path. If they align to one switch or VLAN, investigate network reachability. If hashrate is present but shares are rejected, inspect pool and firmware configuration before pulling machines. If several miners show the same board-level symptom after a maintenance event, suspect process or parts quality before calling it random failure.

Preserve the baseline before making changes. Capture affected miner IDs, board and chip health, logs, recent configuration events, breaker readings, temperatures, and pool-side results. That evidence prevents a familiar failure mode: a team fixes the visible symptom, then cannot explain the cause to customers or prevent recurrence.

3. Contain access and configuration risk

Any unexplained worker, pool, wallet, firmware, or remote-access change should trigger a control-plane check. Freeze nonessential bulk actions. Verify who made the last change, which credentials were used, whether the change touched one group or many, and whether hashrate is being credited where it should be.

Do not assume a pool issue is only a pool issue. A compromised credential, insecure remote access path, or copied configuration template can affect multiple sites. Revert only after you understand the scope. Blindly pushing an old template can overwrite customer-specific settings or reintroduce a known bad endpoint.

4. Communicate facts, not optimism

For high and critical incidents, send an initial internal or customer-facing update early: what is affected, what is not affected, what actions are underway, and when the next update will arrive. Avoid invented restoration times. Operators trust a precise statement such as “216 miners on Container 4 are isolated after a breaker event; no other containers are affected” more than “we are working on it.”

Hosting businesses should tie the incident record to affected client inventory, estimated lost hashrate, SLA terms, and settlement evidence. The operational team needs room to repair the fault, while finance and client success need a clean audit trail for credits and disputed invoices.

Diagnose From Site to Chip

A good incident process does not force every fault into the same checklist. It gives the team a repeatable order of operations and lets evidence choose the branch.

For electrical incidents, verify upstream availability, breaker state, measured current, phase balance, PDU status, cable condition, and the restart sequence. A breaker that trips only after miners return may be revealing overloaded design assumptions, startup inrush behavior, failed load balancing, or a damaged unit pulling abnormal current. Restart in controlled groups and watch real amperage, not just the number of miners online.

For thermal incidents, compare affected miners against neighboring machines. Check inlet temperature, fan RPM, exhaust path, dust loading, heat recirculation, and frequency or power settings. A miner can appear online while producing less hash and accumulating thermal stress. That is not recovered production. It is a delayed failure waiting for the next hot day.

For hashboard incidents, move from machine symptoms to board and chip evidence. Missing chains, degraded chip counts, unstable frequency, repeated nonce errors, and temperature variance tell different stories. Replace a board when the evidence supports it; do not waste technician time swapping entire miners if a board-level fault is already clear. Conversely, a cluster of similar board faults may point to power quality, firmware settings, or environmental exposure rather than bad boards alone.

For network and pool incidents, separate miner connectivity from pool acceptance. Can the miner resolve and reach the endpoint? Is it submitting shares? Are accepted shares arriving at the intended account? Are rejects elevated across the group? This distinction avoids the expensive mistake of dispatching technicians to machines that are healthy but pointed at the wrong destination.

Restore Production Without Creating a Second Incident

Recovery is a controlled return to service, not the moment a dashboard turns green. Bring equipment back in staged groups based on circuit capacity, cooling headroom, and network stability. Watch breaker loads, temperatures, fan behavior, hashrate, reject rates, and pool-side crediting through the first operating window.

Set explicit exit criteria before declaring the incident resolved. The affected group should be online, operating within electrical and thermal limits, submitting accepted shares to the correct account, and free of the alarm pattern that triggered the response. If the incident involved security or customer billing, confirm that access controls and accounting evidence are also restored.

This is where unified operations matter. MinersMe Cloud can connect breaker load monitoring, board and chip diagnostics, pool controls, maintenance tickets, and billing evidence in one operational record. The value is not another dashboard. It is eliminating the gap between seeing a failing board, assigning the repair, confirming the miner’s recovery, and proving the customer’s production impact.

Turn Every Incident Into a Better Operating Standard

Close the incident with a short, blunt review while the facts are fresh. Identify the trigger, root cause or best-supported cause, detection gap, containment action, recovery action, customer impact, and owner for each corrective task. “Technician error” is not a root cause. Ask what control, checklist, permission boundary, alert threshold, training gap, or equipment condition made the error possible.

Measure the time from fault onset to detection, detection to assignment, assignment to containment, and containment to stable production. These intervals expose where the operation is actually slow. A fast technician cannot compensate for alerts buried in email, unclear ownership, missing spare boards, or a maintenance queue managed through scattered messages.

The best mining incident response program becomes quieter over time. Fewer breaker trips become surprise events because load trends identify risk early. Fewer board failures become emergencies because degradation triggers planned repair. Fewer customer disputes become arguments because every outage, action, and recovery has evidence behind it. Build that discipline one incident at a time, and the next bad night becomes a contained operational task instead of a fleet-wide loss.

See it on your own fleet: create a free account, install the agent, or open the live demo — full fleet-to-chip monitoring is included in Pro at $0.40/miner.

More from the blog