+1 (740) 926-6856

Monitoring Cassandra: The Handful of Metrics That Actually Predict Pages

Most clusters we review are already exporting several hundred metrics per node. Almost none of them are on a dashboard anyone reads, and the alert that finally fires is usually client-side p99 latency — which is a lagging indicator. By the time read latency doubles, the cause is hours old.

This is a shorter list. These are the signals that lead an incident, the thresholds we start from, and the virtual tables that let you answer questions from a cqlsh prompt at 3am when Grafana is one more thing to log into.

The four leading indicators

1. Pending compactions

org.apache.cassandra.metrics:type=Compaction,name=PendingTasks, or nodetool compactionstats.

This is the single best early warning Cassandra emits. A healthy node sits near zero and spikes briefly after flushes or repair streams. A node that holds a double-digit backlog for hours is falling behind: SSTable counts climb, reads touch more files, and read latency degrades later, after the backlog has been growing for a while.

What we use in practice:

  • Warn at a sustained backlog above 20 for 30 minutes.
  • Page at a sustained backlog above 100, or any backlog that has grown monotonically for two hours.

The common causes, in the order we check them: compaction throughput throttled too low (nodetool getcompactionthroughput), too few compactor threads for the write rate, a repair that dumped a large number of small SSTables onto the node, or a compaction strategy mismatched to the write pattern. Raising throughput is the fastest lever and also the one that steals disk bandwidth from reads — do it deliberately, one node at a time, and watch read latency while you do.

2. SSTable count per table

nodetool tablestats <keyspace>.<table>, or the LiveSSTableCount metric.

Reads in Cassandra are cheap when a partition lives in a small number of SSTables and expensive when it is smeared across many. Bloom filters and the partition index keep this from being linear, but the relationship is real and it is the mechanism by which a compaction backlog becomes a latency incident.

Watch the per-table number, not the node total. A table under Unified Compaction Strategy or Leveled Compaction running steadily in the low tens is fine. The same table at 300 SSTables means compaction has given up on it. The paired metric worth graphing is SSTablesPerReadHistogram p99 — if that number drifts above roughly 5–8, your read path is doing real work per query.

3. Dropped mutations and hints

nodetool tpstats (or nodetool tablestats for droppable data), plus the DroppedMessage metrics by verb.

A dropped mutation is a write that a replica accepted into its queue and then discarded because it exceeded write_request_timeout_in_ms before it was processed. The coordinator may still have satisfied LOCAL_QUORUM from the other replicas, so the client saw success. The cluster is now inconsistent, and the only thing that will fix it is a hint replay or a repair.

This is why dropped mutations matter more than their raw count suggests: any non-zero rate is a consistency event, not just a performance one. We alert on any dropped MUTATION or READ_REPAIR message at all, then triage on volume. Sustained drops on one node usually mean that node is resource-starved — GC pauses, disk saturation, or a noisy neighbor. Drops across every node usually mean the cluster is genuinely over its write capacity.

Pair it with hints: nodetool statushandoff and the hints-delivered metrics. Growing hint backlog means a replica is unreachable or too slow to keep up. Hints expire (default three hours); past that window, only repair reconciles the divergence, and gc_grace_seconds is quietly counting down.

4. Coordinator vs. local read latency

Cassandra exposes both ClientRequest latency (what the coordinator saw, including the network hop to replicas) and Table-level local read latency (what this node's storage engine did). Graph them together.

The gap between them is the diagnosis:

  • Local latency high, coordinator gap small — storage-engine problem. Too many SSTables, wide partitions, tombstones, disk saturation.
  • Local latency fine, coordinator latency high — the coordinator is waiting on a slow replica, on cross-DC traffic, or on GC pauses somewhere in the request fan-out. Speculative execution in the driver hides some of this; it does not fix it.

Always read p99 and p999, never means. Cassandra's latency distribution is long-tailed by construction, and an average will look healthy while a tenth of a percent of requests time out.

The two metrics people over-weight

CPU utilization. Cassandra nodes routinely run hot during compaction and that is the system working correctly. CPU is useful as context for another signal, not as an alert on its own.

Disk usage percentage. Important, but the threshold is not 90%. Compaction needs free space to work — Size-Tiered compaction can transiently require headroom on the order of the size of the tables being merged. We treat 60% as the warning line and 70% as the point where you should already have a capacity plan, not a ticket.

Virtual tables: answering questions without leaving cqlsh

Cassandra 4.0 added virtual tables in the system_views keyspace, and they are underused. During an incident they let you query node state with CQL instead of shelling into every host for nodetool output.

The ones we reach for:

SELECT * FROM system_views.clients;

Who is connected, from what address, which driver version, and what protocol. Invaluable when an application deploy is the actual cause and nobody has said so yet.

SELECT keyspace_name, table_name, memtable_live_data_size, local_read_latency
  FROM system_views.local_read_latency;
SELECT * FROM system_views.thread_pools WHERE name = 'CompactionExecutor';

Pending and blocked task counts per pool — the same data as tpstats, queryable.

SELECT * FROM system_views.sstable_tasks;

Every running compaction, its progress, and its unit. This is the table that tells you whether a backlog is moving or stuck.

SELECT * FROM system_views.settings WHERE name = 'compaction_throughput_mb_per_sec';

The running configuration as the node actually loaded it — not as cassandra.yaml claims on disk, which matters more often than you would like.

Virtual tables are per-node: you get the state of the coordinator you happen to be connected to. That is a feature during triage (connect to the suspect node directly) and a trap if you assume the answer is cluster-wide.

Getting the metrics out

Cassandra publishes metrics over JMX. The two common paths are a JMX exporter sidecar scraped by Prometheus, or the Metric Collector for Apache Cassandra (MCAC), which bundles a collector plus reasonably sane Grafana dashboards. On Kubernetes, cass-operator wires up metrics endpoints for you.

Whatever you pick, two rules save pain later:

  1. Keep per-table cardinality under control. A cluster with a thousand tables and full per-table metric export will overwhelm a modest Prometheus instance. Export node-level and thread-pool metrics for everything; restrict per-table histograms to the tables that carry your traffic.
  2. Retain at least 30 days. Capacity conversations, compaction-strategy changes, and upgrade comparisons all need a baseline from before the change. Two weeks of retention is how teams end up arguing from memory.

A dashboard we would actually keep

One screen, eight panels:

  1. Coordinator read/write p99 and p999, by datacenter.
  2. Local read latency p99, top five tables by volume.
  3. Pending compactions per node.
  4. Live SSTable count, top five tables.
  5. Dropped messages per node, by verb.
  6. Hints in progress and hints delivered.
  7. GC pause time, p99 and max, per node.
  8. Disk used percentage per node, with the 60% line drawn.

Everything else is a drill-down. If a panel has never been looked at during an incident, it belongs on a second page.

The honest caveat

Monitoring does not fix a data model. A table with 2 GB partitions or a read pattern that requires ALLOW FILTERING will produce alerts forever, and no threshold tuning changes that. Metrics tell you where the cluster hurts; the schema usually tells you why. When the same three tables own every alert on the board, the answer is a modeling change, not another dashboard.


If your dashboards are noisy and your pages still arrive without warning, our cluster and data-model review covers monitoring posture alongside schema, compaction, and repair. For ongoing coverage — alert tuning, hygiene, and escalation — see the operations retainer, or just tell us what your cluster is doing.