3 min readRishi

Read-Your-Writes: The User Just Saved and the Read Replica Does Not Know

Read-Your-Writes: The User Just Saved and the Read Replica Does Not Know

The user edits their display name, the API returns 200, the next GET shows the old name. Refresh again and it is correct. The write went to the primary. The read went to a replica that had not applied the change yet. Replication lag is usually small and, for this one request, it was large enough to be visible. That is read-your-writes, broken.

Replicas exist so reads do not all hit the primary. Sending every read to the primary "to be safe" gives the lag back as load. The requirement is narrower. This user, for a short time after their write, must read a version at least as new as what they wrote. Everyone else can be a second behind.

Ways to pin the next read

ApproachWhat you trackCost
Read-after-write to the primaryNothing. The session's next read, or the next N seconds, uses the primarySimple. A chatty client can pin too much traffic
Replica only if it has caught upThe primary's log position at commit time, compared to the replica's positionNeeds a position the client or the gateway can pass
Cache the user's writeThe new row in a cache keyed by user, read before the replicaExtra store. Must expire or you serve the cache forever
Monotonic tokenA version in the response. The next request sends it. The router picks a node that has that versionClean for apps that already thread a token

The token version is the one that scales. The write response includes the position. The client sends it on the following read. A replica that is behind that position is skipped. If none are caught up, the read hits the primary. Lag becomes a routing decision instead of a user-visible lie.

What people ship by accident

A load balancer that spreads reads at random, with no stickiness and no version, will fail this on every write that is followed quickly by a read. The UI pattern "save, then refetch" is exactly that sequence. Optimistic UI hides it until the refetch overwrites the optimistic value with the stale one, which looks like the save undid itself.

Sticky sessions by user id send that user to one replica. They do not help if the write went to the primary and the sticky node is a different replica. Stickiness without a version is the wrong tool.

Read-your-writes is not "everyone sees the latest." Another user can still observe the old name for the lag window. That is a weaker guarantee, and it is the one replicas can keep. If the product needs every user to see the new name at the same moment, you are asking for something closer to linearizability, and a replica read will not give it to you. Say which one the screen needs before you pick a database flag.

A test you can run without a framework

Write a row. Immediately read it through the path the app uses, fifty times in parallel from one session, and count how often the old value comes back. Then do it from a second session that did not write. The first number is your read-your-writes bug. The second is ordinary lag, which you may choose to allow. If the first number is not zero, the save button is lying, and no amount of client-side cache will fix the refetch that replaces good local state with a stale replica.

Keep reading

Newsletter

New posts, straight to your inbox

One email per post. No spam, no tracking pixels, unsubscribe anytime.

Comments

  • No comments yet. Be the first.