← Blog · Guides & insights · August 20, 2026
Mining Operations Checklist for ASIC Fleet Uptime
A miner does not need to be fully offline to be costing money. A weak hashboard, a rising intake temperature, a breaker nearing its continuous-load limit, or a worker pointed at the wrong pool can drain production for hours before anyone sees a red status light. This mining operations checklist is built for the work that protects real fleet revenue: finding the conditions that become outages before they spread through a container or site.
The point is not to create another spreadsheet that technicians update after the problem is over. The point is to establish a repeatable operating rhythm, assign ownership, and connect every alert to a response. A 200-miner site can get away with some manual review. A hosting operation with thousands of units cannot. At scale, the checklist must become an automated workflow backed by live telemetry, maintenance records, and clear escalation rules.
Mining Operations Checklist: Start With Fleet-Level Reality
Start each shift at the fleet view, not at an individual miner. Compare online hashrate against expected hashrate by site, building, container, customer, miner model, and firmware group. The fleet total can look acceptable while one customer allocation, one electrical panel, or one rack is quietly underperforming.
Review offline, zero-hash, degraded, and unstable miners separately. These are different failure states and require different action. A zero-hash miner may need a fast reboot or board inspection. A miner cycling online and offline may indicate thermal protection, an unstable power supply, poor network conditions, or a firmware issue. A unit hashing below target may still be producing, but it can also be the first signal of chip degradation or an intake-air problem.
The operating lead should also review hashrate variance against the prior shift and the prior day. A sudden site-wide decline is rarely a collection of unrelated failed machines. Look for shared causes: a pool endpoint issue, DNS failure, switch fault, transformer event, cooling degradation, or an electrical limit being reached.
Verify Pool, Worker, and Wallet Integrity
Pool configuration is production infrastructure. Treat it with the same seriousness as a main breaker. Confirm that each customer or internal fleet segment is hashing to the approved pool, account, worker naming convention, and payout destination. A miner can be electrically healthy and still send revenue somewhere it does not belong.
Check rejected-share rates, stale-share rates, connection failures, and unexpected fallback-pool use. A short burst of stale shares may be harmless during a network event. A persistent rise points to latency, packet loss, overloaded networking equipment, bad pool routing, or a pool-side problem. The response depends on whether the issue is isolated to a miner group or visible across an entire site.
Watch for unauthorized worker changes and unfamiliar wallet addresses. Do not leave these for a weekly audit. Pool misconfiguration and stolen worker hashpower are incidents. Preserve the affected miner list, identify when the configuration changed, restore approved settings, and review who had remote access.
Check Electrical Load Before the Breaker Makes the Decision
Electrical visibility has to be granular enough to show what is happening before a trip. Review live amperage by breaker, PDU, panel, and phase where instrumentation allows it. Compare measured load to the continuous operating limit, not merely the nameplate rating. A breaker that trips during peak heat or a restart event was already operating too close to the edge.
Look for imbalance between phases, rising load on a circuit after equipment changes, and unusual differences between similar containers. Also compare electrical consumption with fleet hashrate. If power rises while hashrate stays flat or falls, the fleet may be wasting energy through degraded hardware, poor cooling, unstable firmware, or a growing share of underperforming units.
When a breaker approaches its defined threshold, the response should be predetermined: identify the connected miners, reduce load in a controlled order, validate readings, and create an electrical ticket. Randomly power-cycling machines after a trip is not recovery. It is a way to turn one fault into damaged equipment and another outage.
Read Temperature as a Failure Forecast
Ambient temperature alone is not enough. Review intake temperature, exhaust temperature, board temperature, chip temperature, fan speed, and thermal error events. A container can show an acceptable average while a row near a blocked intake or a failed extraction fan runs far hotter than the rest.
Focus on drift. A board that runs several degrees hotter than comparable boards in the same model and environment is more valuable as an early maintenance target than as a future emergency. Likewise, a fan operating at unusually high speed may be compensating for dust buildup, restricted airflow, or a weakening cooling path.
Technicians should physically inspect recurring thermal zones. Check filters, louvers, fan walls, cable obstructions, hot-air recirculation, and damaged containment. In immersion environments, review fluid temperature, flow, pump behavior, and signs of leaks or contamination. The right threshold depends on miner model, firmware, site design, and climate. The principle does not change: thermal anomalies need a work order before they become board failures.
Diagnose From Miner to Board to Chip
A useful monitoring system lets operators descend from a weak fleet segment to a specific miner, then to the hashboard and chip behavior causing the loss. Review board-level hashrate, chip counts, voltage readings, ASIC errors, frequency behavior, and board temperature deltas. A miner that reports as online may have one board producing a fraction of its expected output.
Do not send every degraded miner directly to the repair bench. First separate recoverable conditions from hardware faults. A controlled reboot, configuration correction, fan replacement, cable reseat, or power-supply check may restore a unit quickly. Repeated board faults, missing chips, abnormal voltage behavior, or persistent low hashrate after standard recovery should move to repair with diagnostic evidence attached.
Each ticket should state the symptom, telemetry at the time of failure, actions already attempted, required parts, assigned technician, and verification result after return to service. “Miner fixed” is not a closure standard. The unit should be observed long enough to confirm stable hashrate, normal temperature, and correct pool assignment.
Run Maintenance by Priority, Not by Noise
A noisy alert feed trains people to ignore alarms. Set priorities based on revenue exposure, safety risk, failure propagation, and customer impact. A breaker overload risk, widespread pool misconfiguration, active thermal shutdown, or suspected wallet change is immediate. A single low-performing board is urgent when repair capacity is available, but it should not distract the team from a site-wide event.
For planned maintenance, review recurring faults by model, location, technician action, and time to recurrence. If the same racks repeatedly show fan failures or board temperature alarms, the root cause may be environmental or electrical rather than individual miners. The repair count is not the metric that matters. The metric is whether the failure returns.
This is where disconnected tools fail operators. If alerts live in one system, tickets in another, breaker readings in a third, and customer credits in a spreadsheet, the team spends its shift reconstructing the incident. MinersMe Cloud is designed to keep fleet telemetry, diagnostics, electrical conditions, maintenance workflows, and commercial consequences in one operating console.
Protect Hosting Commitments and Settlement Data
For hosting providers, the checklist does not stop at miner health. Confirm customer fleet assignments, uptime calculations, billing periods, energy rates, repair charges, SLA rules, and credit triggers. A technically correct repair can still become a customer dispute if the outage window, machine ownership, or charge is not documented.
Review exceptions before invoices go out. Look for miners moved between customer groups, prolonged repair status, missing meter data, manual rate overrides, and unresolved downtime events. When the data trail is clean, client-success teams can explain a credit or charge with facts instead of chasing technicians through chat logs.
End Every Shift With an Owned Next Action
The final pass is simple: every abnormal condition must be resolved, scheduled, or explicitly handed off. Open tickets need an owner and a due time. Pending parts need an expected arrival and a list of affected miners. Site-level risks need a named decision-maker, not a note buried in a shift report.
A disciplined mining operation is not one that never sees failures. ASIC fleets run hot, consume heavy power, and expose every weak connection in the facility. The operation that wins is the one that sees the failing board, rising breaker load, bad worker, or heat pattern early enough to act while the outage is still preventable.
See it on your own fleet: create a free account, install the agent, or open the live demo — full fleet-to-chip monitoring is included in Pro at $0.40/miner.
MinersMe Cloud