Skip to main content

MVCC and Transactions

Kahuna's key/value transaction engine combines MVCC, write intents, locks, operation registration, finalization fencing, two-phase commit, and optional durable commit decisions. This is what lets scripts and interactive transactions update multiple keys atomically across partitions.

Transaction Coordinator​

TransactionCoordinator owns transaction sessions and coordinates distributed reads, writes, validation, finalization, and cleanup. Kahuna Script execution uses this transaction machinery, and interactive client sessions use the same coordinator-owned lifecycle.

The important invariant is that the server owns the transaction working set. Clients may carry a transaction handle, but they do not send a final list of touched keys to commit. As commands succeed, the coordinator records confirmed effects and later commits or rolls back from that server-owned state.

It can optimize simple one-command scripts by dispatching directly to the relevant command. Multi-statement scripts, explicit transaction blocks, and interactive sessions use the full transaction path.

Transaction IDs​

Transactions use Hybrid Logical Clock timestamps as transaction IDs. HLC timestamps provide ordering that works across nodes while still preserving causal progression.

A transaction context tracks:

  • Transaction ID
  • Coordinator key and optional durable record anchor
  • Modified keys and durability
  • Point, prefix, and range locks still held
  • Latest-read observations and validation policy
  • Snapshot timestamp policy
  • Registered operations and pending operation count
  • Variables and script parameters for script execution
  • Transaction status
  • Locking mode
  • Admission priority
  • Timeout
  • Decision durability mode
  • Finalization state and frozen working-set snapshot

Interactive operations also carry stable operation IDs. The operation registry lets Kahuna identify duplicate retries, reject reused IDs with different inputs, and avoid applying the same mutation twice after a lost response.

MVCC Entries​

KeyValueEntry stores the current value and metadata for a key. It can also hold MvccEntries, keyed by transaction ID. Each KeyValueMvccEntry contains a proposed or versioned value plus revision, expiration, state, and HLC metadata.

This lets Kahuna keep transactional state separate from the committed current value until the transaction commits.

The first latest transactional point read pins a committed value or absence per key, including a committed unsettled intent. These pins do not form a common transaction-start snapshot. Follow-up point reads or writes can abort immediately if the observation is stale; finalization still validates registered dependencies.

Snapshot reads are different. A read at a fixed HLC timestamp selects the newest revision whose commit timestamp is at or before that timestamp. Because it is a historical view, it does not create a live read dependency for latest-state validation.

Write Intents​

A write intent records that a transaction intends to modify a key. Kahuna uses write intents to prevent conflicting updates and to make commit/rollback deterministic.

There are two important levels:

  • Key write intent: protects a single key.
  • Prefix write intent: protects a bucket or prefix, which matters for operations such as get by bucket.
  • Range write intent or lock: protects an ordered interval, which matters for bounded range reads.

Prefix and range locks block conflicting mutations while held. They are leader-local and can disappear on failover; range renewal is best-effort, so validation remains necessary.

Locking Modes​

Kahuna supports two transaction locking styles:

ModeBehavior
PessimisticAcquire locks before or during operations. Point reads can use point locks, bucket reads can use prefix locks, and range reads can use range locks.
OptimisticRead without exclusive locks, then validate read observations and concurrent write intents at commit time.

Pessimistic locking reduces conflict surprises but can hold locks longer. Optimistic locking allows more read concurrency but requires conflict detection during prepare.

Operation Registration and Working-Set Folding​

Interactive transaction operations follow a registration path:

BeginOperation
-> participant execution
-> CompleteOperation with confirmed effects
-> fold effects into TransactionContext

Only confirmed effects are folded into the working set:

  • Successful writes, deletes, and extends record modified keys.
  • Successful point, prefix, and range lock acquisitions record held locks.
  • Explicit releases remove matching lock descriptors.
  • Latest point reads record read observations for cleanup and optional validation.
  • Snapshot reads do not record live read dependencies.

Failed conditional writes and failed lock acquisitions do not enter the working set. This is what lets finalization avoid relying on client-side bookkeeping.

Finalization Fence​

Commit, rollback, close, and abandoned-session cleanup share one finalization slot. The first finalizer closes the transaction to new operations, waits for operations registered before the fence to drain, freezes the working set, and then runs commit or rollback from that immutable snapshot.

A retryable finalization failure can release the attempt slot for another commit or rollback call, but it never reopens the session to new reads or writes.

Two-Phase Commit​

For multi-key transactions, Kahuna uses a prepare/commit protocol. The concrete path depends on the frozen working set:

Working setCommit path
Read-onlyNo prepare round is needed.
All ephemeralCommit in memory.
All persistentDurable-intent two-phase commit.
Mixed under BestEffortPrepare ephemeral mutations first, then let the persistent decision drive ephemeral commit or rollback. The ephemeral subset can be lost on process failure.

The general lifecycle is:

  1. Start the transaction and assign a transaction ID.
  2. Read and write through participant leaders.
  3. Stage writes as MVCC entries and write intents.
  4. Fold confirmed effects into the coordinator working set.
  5. Close the transaction to new operations before finalization.
  6. Prepare mutations on all participant partitions.
  7. Validate reads when the policy requires it.
  8. Decide commit or abort.
  9. Release locks and clean up transaction state.

For persistent writes, the canonical commit decision determines visibility across prepared participants; physical settlement can lag that decision. Explicit Durable mode rejects ephemeral modified keys. Mixed BestEffort transactions do not promise crash atomicity for the ephemeral subset.

Read-only transactions can commit without a prepare round.

Durable Commit Decisions​

Live session outcomes are retained in memory for a bounded idempotency window. Persistent modifications use durable-intent finalization even under the default BestEffort policy; explicit Durable mode additionally rejects ephemeral modified keys.

The first confirmed persistent modified key becomes the record anchor. The coordinator initializes a canonical transaction record on that anchor partition, prepares the anchor partition's intents in the same ordered proposal when possible, replicates prepared intents for every other modified persistent partition, validates staged bases and reads, then compare-and-sets the canonical record to Commit or Abort.

Persistent participants store completion receipts when committed values are applied. Recovery uses those receipts to distinguish "already committed" from "unknown" after original intents are gone.

The durable commit ordering is:

prepare durable intents
-> decide canonical record
-> return Committed when deferred settlement is enabled
-> materialize committed values with receipts
-> settle prepared intents

Durable decision mode does not persist the active interactive session. If the coordinator disappears before a canonical record is installed, the session is lost like a best-effort transaction. Once prepared intents exist, participant leader changes do not lose the staged value; recovery can resolve the intents from the canonical record. Ephemeral modified keys are rejected in durable mode because their values, intents, and receipts cannot survive process loss.

With default deferred settlement, a committed transaction can return before materialization and settlement finish. While a committed intent is still pending, point reads, scans, and writes consult the canonical record, locally or by routing to the anchor leader, and resolve the intent without serving the stale previous value.

Durable materialization is by reference by default. The committed log record names the prepared intent and carries the canonical timestamp and revision, while each replica reads the value from its local prepared-intent store. This keeps MVCC visibility correct and avoids copying the value through the Raft log a second time.

Unprepared session-owned write intents and no-expiry range locks have a liveness ceiling. If the owning transaction disappeared and the ceiling has passed, the next touch can drop the orphaned state instead of letting it block reads or scans forever. Prepared durable intents are exempt because their outcome belongs to the canonical decision record.

For the end-to-end path, see Transaction Lifecycle.

Revisions and Snapshots​

Each key tracks a revision counter. Reads can request a specific revision, and recent revisions can be cached. Transactions use revision and HLC metadata to decide which value is visible and whether a concurrent write invalidates a commit attempt.

SET ... NOREV still advances the current revision but intentionally skips the archived historical revision record. That reduces write amplification for cache-style values, but it also means a skipped revision cannot be served later by GET ... AT <revision> or by a historical snapshot read that needs that exact archived version.

Snapshot Floor​

The MVCC snapshot floor protects historical reads that must remain valid for longer than the normal revision-retention window. A client acquires a leased snapshot hold at timestamp T; while that hold is live, cleanup must keep the revision that was visible at or before T, plus every newer revision.

The reported effective floor is the minimum timestamp across live holds. The pruning floor includes all registered holds, even expired ones, until replicated release or purge. Hold state is replicated through the system partition and restored from local snapshots. Acquire, renew, and release operations can enter through any node, but they are routed to the system-partition leader before being committed.

The floor constrains both revision cleanup paths:

  • In memory, Kahuna keeps the normal newest RevisionRetention revisions, the floor boundary, and committed archive revisions whose backend flush is still pending. Unflushed revisions can exceed the retention target.
  • On disk, persistent revision cleanup must not delete the boundary revision or anything newer, even when count-based or age-based retention would otherwise remove it.

Historical reads first try the in-memory archive. If the requested timestamp is older than the in-memory window, persistent read paths fall back to on-disk revision history. Point reads, range reads, bucket reads, and prefix scans all use the same rule: return the newest revision whose commit timestamp is at or before the requested snapshot timestamp.

When a snapshot floor pins a boundary revision, the in-memory archive can become the pinned boundary plus the newest retained revisions. If the revisions between them have been trimmed from memory, Kahuna treats a boundary hit inside that gap as a cache miss and falls back to persisted revision history. This prevents a held floor from returning an older value when a newer disk-only revision is actually visible at the requested timestamp.

Large snapshot scans can continue through backend revision history without loading every scanned key back into the hot actor cache. This keeps historical analytics and recovery-style reads from displacing current working-set entries just because old revisions live on disk.

Persistent cleanup is budgeted. A targeted cleanup pass deletes at most PersistentRevisionCleanupBatchSize records and stops when PersistentRevisionCleanupTimeBudget expires, leaving remaining keys queued for later. This keeps revision pruning from monopolizing the persistence writer when a small set of keys has very deep history.

After restart, durable snapshot holds are loaded with a startup grace window before expired leases are purged. That grace keeps a timestamp protected long enough for a holder to renew after full-cluster downtime, provided the underlying historical revisions are still present.

The hold API is described in Snapshot Holds.

Historical Read Fences​

Safe-time waits cover live writers that may commit at or before the requested timestamp. The serving actor folds eligible read timestamps into its HLC, and durable commit timestamps are minted above the maximum participant staged timestamp. A read more than five seconds ahead of the serving HLC skips the clock fence; kahuna.kv.snapshot_clock_fence_skipped_total counts this path. Such future reads can change as later commits land within the requested time.

Disk-history fallback returns MustRetry while needed committed revisions are unflushed. kahuna.kv.revisions.retained_unflushed_total counts retention beyond the memory target; kahuna.kv.revisions.history_reads_fenced_total counts fenced fallback reads. TTL uses the current read time, not historical time. Snapshot holds protect retention and cannot recreate pruned or NOREV history.

See Transaction Reads and Locks for newer-head precedence over lingering intents, point-lock convergence, and range-lock compatibility. Durable Settlement covers the optional materializing-resolve encoding.