+1 (740) 926-6856

Securing a Running Cassandra Cluster: Auth, TLS, and Audit Logging Without a Window

Most of the clusters we inherit were built in a hurry. Somebody needed a datastore for telemetry, authenticator: AllowAllAuthenticator was the default, internode traffic was plaintext because TLS "can come later," and five years on the cluster holds customer data under a compliance regime nobody anticipated. The hard part is never deciding to turn security on. The hard part is turning it on without a window of downtime on an always-on cluster.

This is the order we run that work in, and the specific step in each phase where people get paged.

Phase 0: know what you are protecting

Before touching cassandra.yaml, write down three things.

Who connects. Every application, every batch job, every analyst with a cqlsh alias, every sidecar exporter and backup agent. Enabling authentication breaks every connection you did not enumerate, and the one you forget is always a nightly job that fails quietly for a week.

What the blast radius of a restart is. Auth, encryption, and audit changes all require a rolling restart. On a healthy cluster with LOCAL_QUORUM reads and RF=3 per DC, one node down is routine. If your cluster is RF=2, or a DC is already carrying a dead node, fix that first — a rolling restart is not the moment to discover you have no replica headroom.

Where the data boundary actually is. Client-to-node TLS protects traffic crossing an untrusted network. Node-to-node TLS matters most across datacenters and availability zones. Encryption at rest protects disks and snapshots — including the snapshots that get copied to object storage, which is where most real exposure lives.

Phase 1: authentication, without locking yourself out

Cassandra's PasswordAuthenticator stores credentials in the system_auth keyspace. That keyspace's replication is the detail that bites.

The default replication for system_auth on a new cluster is minimal. If the node holding a user's credentials is down, logins fail — including the login your recovery procedure needs. Before enabling auth:

ALTER KEYSPACE system_auth WITH replication =
  {'class': 'NetworkTopologyStrategy', 'dc1': 3, 'dc2': 3};

Then run a full repair of system_auth. Replication changes do not move data; repair does. Skipping this is the single most common way teams lock themselves out of a cluster they just secured.

The rollout itself is a rolling change — set authenticator: PasswordAuthenticator and authorizer: CassandraAuthorizer in cassandra.yaml, restart one node at a time. During the roll the cluster is mixed: nodes already restarted demand credentials, nodes not yet restarted do not. That is fine for drivers that have credentials configured (they are simply ignored by the un-migrated nodes), which means the correct sequence is deploy credentials to every client first, then roll the cluster. Clients that present a username and password to an AllowAllAuthenticator node connect normally. Clients that present nothing to a PasswordAuthenticator node do not.

Immediately after the roll: change the default cassandra/cassandra superuser password, create a separate named superuser, and disable or lock the default account. The default superuser also reads at QUORUM rather than LOCAL_QUORUM — another reason not to leave it as your operational identity in a multi-DC cluster.

Also raise the auth caches (credentials_validity_in_ms, permissions_validity_in_ms, roles_validity_in_ms) from their defaults if you see a latency bump on connection-heavy workloads. Short validity windows mean every new connection hits system_auth.

Phase 2: roles that match how teams actually work

CassandraAuthorizer gives you GRANT/REVOKE over keyspaces, tables, and functions, with roles that can be granted to other roles. Use that nesting; do not create one role per human.

A structure that survives audits without creating daily friction:

  • One role per application, granted SELECT/MODIFY on exactly the keyspaces it owns. Applications get no ALTER, no DROP, no cluster-wide DESCRIBE.
  • A read-only role for dashboards, exporters, and support queries.
  • A schema-change role used only by migration tooling in CI, not by people.
  • An operator role for humans, with permissions broad enough to debug and narrow enough that DROP KEYSPACE is a deliberate escalation, not a typo away.

Grant per keyspace, not per table, unless you have a real reason — table-level grants drift the moment someone adds a table and forgets the GRANT, and the failure surfaces as an application error in production at 2 a.m.

Phase 3: TLS, in the order that avoids a flag day

Both internode and client encryption support an intermediate mode, and those modes are the entire trick.

Internode first. Set server_encryption_options.internode_encryption: all with optional: true and roll. In this state a node accepts both encrypted and plaintext internode connections, so a half-rolled cluster stays fully connected. When every node has restarted and nodetool netstats plus your logs confirm the cluster is gossiping over TLS, set optional: false and roll again to refuse plaintext. Two rolls, zero downtime, no moment where the ring is split.

Clients second. client_encryption_options.enabled: true with optional: true makes the node listen for both TLS and plaintext CQL on the native port. Roll it, then migrate application drivers one deployment at a time, confirming each is connecting over TLS (driver-side connection logs, or the system_views/nodetool clientstats view of connections). Once nothing plaintext remains, set optional: false and roll a final time.

If you are also doing client certificate auth (require_client_auth: true), stage it after plain TLS is universally working. Combining "turn on encryption" and "turn on mutual auth" in one change means every failure is ambiguous.

Plan certificate rotation at the same time you plan the rollout. A cluster whose internode certificates all expire on the same day in two years is a scheduled outage you have already booked. Stagger expiries, document the rotation runbook, and alert on days-to-expiry — not on handshake failures, which is the alert that fires after the outage starts.

Phase 4: encryption at rest, and the snapshot gap

Transparent data encryption is a DSE feature; open-source Cassandra does not encrypt SSTables itself. On OSS clusters the practical answer is volume-level encryption — LUKS, EBS encryption, or the cloud provider's equivalent — which is cheap, nearly free in CPU terms on modern hardware, and covers the actual threat (a disk or volume snapshot leaving your control).

The gap people miss is backups. Snapshots copied to object storage are full copies of your data outside the encrypted volume. Encrypt the bucket, scope the credentials the backup agent uses to write-only where possible, and verify the restore path still works with encryption in place — a restore procedure that nobody has exercised since you added encryption is not a restore procedure.

Commit logs and hints also hold recent writes in plaintext on disk. If your volume encryption covers only the data directory and not the commitlog/hints directories, you have left the most recent data the least protected.

Phase 5: audit logging that someone will actually read

Cassandra 4.0 added native audit logging (audit_logging_options), writing to a binary log that auditlogviewer reads. It is genuinely useful and genuinely capable of hurting you if configured carelessly.

Audit everything and you will write more log volume than data, and the first symptom is disk pressure on the node whose audit directory shares a volume with the data directory. Put audit logs on their own volume, set roll_cycle, and cap the archive.

The configuration that earns its cost: include the DCL and DDL categories always (role changes, grants, schema changes — the events that matter for an audit trail and are rare enough to be free), include authentication events (especially failures), and either exclude DML entirely or restrict it to the specific keyspaces where a compliance regime requires statement-level records. included_keyspaces and excluded_categories do this work. Excluding your own operational roles from DML audit is reasonable; excluding them from DCL audit is not.

Full query logging (nodetool enablefullquerylog) is the adjacent tool, and it is for debugging and replay, not for audit. Turn it on for a bounded window, pull the capture, turn it off.

The sequence, compressed

Enumerate clients → fix system_auth replication and repair it → push credentials to clients → roll auth and authorizer → build roles, rotate the default superuser → internode TLS optional, then strict → client TLS optional, migrate drivers, then strict → volume encryption including commitlog, hints, and backup buckets → audit logging scoped to DCL, DDL, and auth events.

Each arrow is a rolling restart or a deployment, and each one is reversible on its own. The teams that have a bad time are the ones that batch three phases into one change window to "save restarts." Restarts are cheap on a healthy cluster. Ambiguous failures at 2 a.m. are not.

If you are facing this work on a cluster that is already carrying production traffic — or you have been handed an audit finding with a date on it — that is a normal shape for our cluster review and operations engagements. Tell us what the cluster looks like and what the deadline is, and we will tell you honestly how many change windows it takes.