Start Debugging

Volatile.Read vs Volatile.ReadBarrier in .NET 10

Volatile.Read is an acquire load of one location. Volatile.ReadBarrier, new in .NET 10, is a fence that gives every earlier read acquire semantics. Use Volatile.Read for flags and published references, and ReadBarrier when you need a batch of plain or non-atomic reads to complete before the next memory access, as in a seqlock.

Volatile.Read(ref x) reads one location with acquire semantics: nothing later in your code can move above that read. Volatile.ReadBarrier(), added in .NET 10, reads nothing at all. It is a fence that gives acquire semantics to every read before it, so a whole batch of plain (even non-atomic) reads must finish before any memory access after the barrier. Use Volatile.Read for the common case of a flag, a counter, or a published reference. Reach for ReadBarrier when you need several ordinary reads, or one read that is too large to be atomic, to complete before a re-check. The textbook case is the reader side of a seqlock.

Everything below was measured on .NET 10.0.10 (SDK 10.0.302), C# 14, on an Apple M4 (arm64). The barrier APIs exist in System.Threading.Volatile from .NET 10 onward; on .NET 9 and earlier there is no public equivalent short of Interlocked.MemoryBarrier().

The comparison at a glance

Volatile.Read(ref x)Volatile.ReadBarrier()
Available since.NET Framework 4.5.NET 10
Reads a valueYes, one locationNo
What gets acquire semanticsThat one readAll reads before the call
Blocks later reads and writes from moving upYesYes
Makes the read atomicYes, for supported types (including long/double on 32-bit)No, atomicity is your problem
Works with any T, structs, native memoryNo, fixed set of overloadsYes, it orders whatever reads precede it
arm64 codegen (measured, .NET 10.0.10)ldapur (load-acquire)dmb ishld (load fence)
x64 codegenplain mov, compiler ordering onlyno instruction, compiler ordering only
Typical useflags, double-checked init, published referencesseqlocks, version-validated caches, batched reads

What the two APIs promise

The .NET memory model spec (docs/design/specs/Memory-model.md in dotnet/runtime) lists both under “volatile reads have acquire semantics”, with one telling footnote on the barrier: it “applies to all prior reads”. Acquire means no read or write that comes later in program order may execute ahead of the acquiring read.

With Volatile.Read(ref _version), the acquire is attached to the load of _version and nothing else. Reads that happened before it in program order are not constrained at all. They may still drift below it.

With Volatile.ReadBarrier(), the acquire is attached to every load that precedes the call. The API proposal (dotnet/runtime#98837) calls this a Read-ReadWrite barrier: all preceding reads must complete before any subsequent memory operation. Its counterpart, Volatile.WriteBarrier(), is a ReadWrite-Write barrier: all preceding memory operations complete before any subsequent write.

So the two APIs are not two strengths of the same thing. They answer different questions:

Neither one is a full fence. A ReadBarrier does nothing to stop an earlier write from being reordered with a later read (the store-load case). If you need that, you still need Interlocked.MemoryBarrier() or an Interlocked operation.

What the JIT actually emits

The JIT treats both methods as intrinsics. The source of Volatile.cs is just [Intrinsic] public static void ReadBarrier() => ReadBarrier();, and the importer replaces the call with a memory-barrier node flagged as load-only (PR #107843). To see what that becomes, I compiled a small class with full opts and dumped it with DOTNET_JitDisasm:

// .NET 10.0.10, C# 14
// DOTNET_TieredCompilation=0 DOTNET_JitDisasm='Codegen:*' dotnet vb.dll
sealed class Codegen
{
    private int _x;
    private long _a, _b, _c, _d;

    [MethodImpl(MethodImplOptions.NoInlining)]
    public long AcquireFour() =>
        Volatile.Read(ref _a) + Volatile.Read(ref _b) +
        Volatile.Read(ref _c) + Volatile.Read(ref _d);

    [MethodImpl(MethodImplOptions.NoInlining)]
    public long PlainFourThenBarrier()
    {
        long sum = _a + _b + _c + _d;
        Volatile.ReadBarrier();
        return sum;
    }

    [MethodImpl(MethodImplOptions.NoInlining)]
    public void BarrierThenPlainFour(long v)
    {
        Volatile.WriteBarrier();
        _a = v; _b = v; _c = v; _d = v;
    }
}

On the M4, the interesting instructions were:

; AcquireFour: four separate load-acquire instructions
ldapur  x1, [x0, #0x08]
ldapur  x2, [x0, #0x10]
ldapur  x2, [x0, #0x18]
ldapur  x0, [x0, #0x20]

; PlainFourThenBarrier: two paired loads, then one load fence
ldp     x1, x2, [x0, #0x08]
ldp     x2, x0, [x0, #0x18]
dmb     ishld

; BarrierThenPlainFour: a full fence, then two paired stores
dmb     ish
stp     x1, x1, [x0, #0x08]
stp     x1, x1, [x0, #0x18]

Three things stand out.

First, Volatile.Read compiles to ldapur, an RCpc load-acquire (the RCpc extensions arrived in ARMv8.3 and v8.4), which the M4 supports. Cores without RCpc get the older ldar instead. Either way, there is no separate fence instruction.

Second, the plain reads before ReadBarrier stay plain, so the JIT is free to pair them into ldp (and, for a 32-byte struct copy, into a pair of 128-bit ldp q loads). You lose that freedom with four acquire loads. That is the efficiency argument the proposal made: one fence for N reads instead of N ordered reads.

Third, Volatile.WriteBarrier() is a full dmb ish on arm64, exactly what Interlocked.MemoryBarrier() emits. The JIT has a comment saying it cannot currently emit a store-only barrier better than a full one on arm64, so do not expect WriteBarrier to be cheaper than a full fence there.

On x64, both barriers emit no instruction at all. The PR’s codegen comment is explicit: load-only and store-only barriers “are no-ops on xarch”, because x86’s TSO model already keeps loads ordered with later loads and stores, and stores ordered with earlier memory operations. They still matter on x64, though: they stop the JIT itself from reordering, caching, or eliminating memory accesses across the barrier. I did not have an x64 machine for this run, so the x64 row in the table comes from the JIT source, not a disassembly.

A seqlock: the case ReadBarrier was built for

The runtime itself was the first customer. GenericCache and CastCache in CoreLib used an internal Interlocked.ReadMemoryBarrier() and were switched to Volatile.ReadBarrier() in the same PR. Their comment spells out the pattern: “we must read in this order: version -> [entry parts] -> version”.

That is a seqlock. A single writer bumps a version to an odd number, writes the data, then bumps it to the next even number. Readers read the version, copy the data with ordinary loads, and read the version again. If both reads match and are even, the copy is consistent. The data can be any size: a 32-byte struct is not atomic on any platform, and that is fine, because the version check catches torn copies.

Here is the minimal version, with both barriers in the places they belong:

// .NET 10, C# 14
struct Snapshot { public long A, B, C, D; }

sealed class SeqLockBox
{
    private int _version;          // even = stable, odd = write in progress
    private Snapshot _data;

    // Single writer only.
    public void Write(long n)
    {
        int v = _version;
        _version = v + 1;          // mark "writing" (odd)
        Volatile.WriteBarrier();   // odd version is published before any data write below
        _data.A = n; _data.B = n; _data.C = n; _data.D = n;
        Volatile.Write(ref _version, v + 2); // release: data writes complete before the even version
    }

    public bool TryRead(out Snapshot snapshot)
    {
        int v1 = Volatile.Read(ref _version); // acquire: the data reads below cannot move above this
        snapshot = _data;                     // plain, non-atomic 32-byte copy
        Volatile.ReadBarrier();               // every read above completes before the re-check
        return (v1 & 1) == 0 && _version == v1;
    }
}

Look at how the reader uses both APIs. The first version read is a Volatile.Read, because we need the data reads to stay below it. The data copy is plain. Then ReadBarrier keeps the data reads above the second version read. No single Volatile.Read can express that second constraint, because Volatile.Read only constrains what comes after the location it reads, and here the thing we need to order is what came before.

The writer mirrors it. Volatile.Write on the final even version is a release, so the data writes cannot sink below it. But a release does nothing to stop the data writes from rising above the earlier odd-version store. WriteBarrier covers that side.

Proving each half is necessary

I ran the reader and writer on two threads for five seconds per scenario and counted how many accepted snapshots had A, B, C, D disagreeing. Each scenario removes one piece of the ordering:

// .NET 10, C# 14: the reader variants in the stress test
public bool TryReadAcquireOnly(out Snapshot snapshot)   // no ReadBarrier
{
    int v1 = Volatile.Read(ref _version);
    snapshot = _data;
    return (v1 & 1) == 0 && _version == v1;
}

public bool TryReadBarrierOnly(out Snapshot snapshot)   // no acquire on the first read
{
    int v1 = _version;
    snapshot = _data;
    Volatile.ReadBarrier();
    return (v1 & 1) == 0 && _version == v1;
}

Results on the M4, .NET 10.0.10, Release build, two runs:

ScenarioAccepted snapshots (run 1 / run 2)Torn and accepted (run 1 / run 2)
No ordering at all (plain reads)1,014,876,206 / 1,003,303,309547,804 / 515,135
Volatile.Read only, no ReadBarrier164,358,676 / 152,032,56199 / 357
ReadBarrier only, plain first read34,543,735 / 27,982,88454 / 62
Writer without WriteBarrier, correct reader354,942,744 / 384,287,32466,155,404 / 54,597,163
Both barriers (the code above)62,659,697 / 66,048,7380 / 0

Every half-measure produced torn data that passed validation. The rare ones are the dangerous ones: 99 bad reads out of 164 million is the kind of bug that survives every test run and shows up in production on a Graviton or Ampere box. The missing WriteBarrier was the loudest failure, and the disassembly shows why: the two writer methods compile to identical code apart from the single dmb ish, so every one of those 54+ million tears is the arm64 core making the data stores visible before the odd version store.

On x64 you would very likely see zero tears for most of these rows, because the hardware does not reorder in those directions. That is exactly why these bugs ship. The code is still wrong on x64, since the JIT is allowed to reorder plain accesses, and it becomes visibly wrong the moment it runs on arm64.

The JIT reorders too, not only the CPU

The first version of my stress harness hung forever, and it is worth showing why. The broken reader looped until it saw an even version:

// .NET 10, C# 14: do not do this
public void WaitForEvenBroken()
{
    while ((_version & 1) != 0) { }
}

The JIT compiled that to one load and a branch to itself:

ldr     w0, [x0, #0x08]
and     w0, w0, #1
G_M000_IG03:
cbnz    w0, G_M000_IG03

The load of _version was hoisted out of the loop, which is legal for an ordinary field read with no intervening synchronization. If the first read happened to land on an odd version, the thread spins forever. A Volatile.Read(ref _version) inside the condition fixes it, and so would a ReadBarrier inside the loop body. This is the part of “volatile” that x64 developers do experience, and it is why the barriers are not empty calls even where they emit no instruction.

When to pick Volatile.Read

In all of these, the ordering is anchored to a single read, so Volatile.Read says exactly what you mean and generates no fence on either architecture.

When to pick Volatile.ReadBarrier

The cost, measured

There are two correct ways to write the seqlock reader without ReadBarrier: make every data read a Volatile.Read, or use a full Interlocked.MemoryBarrier() where the barrier goes. I benchmarked all of them against the unordered (broken) reader with BenchmarkDotNet 0.15.8. Each invocation does 1,024 single-threaded reads of the 32-byte snapshot, and the table reports the cost per read:

// .NET 10.0.10, C# 14, BenchmarkDotNet 0.15.8, Apple M4 (arm64)
[Benchmark(OperationsPerInvoke = N)]
public long VolatileReadPlusReadBarrier()
{
    long sum = 0;
    for (int i = 0; i < N; i++)
    {
        int v1 = Volatile.Read(ref _version);
        Snapshot s = _data;
        Volatile.ReadBarrier();
        if ((v1 & 1) == 0 && _version == v1) sum += s.A + s.B + s.C + s.D;
    }
    return sum;
}
Reader (per snapshot read)MeanRatio
No ordering (broken)0.916 ns1.00
Volatile.Read on the version and on all four fields1.135 ns1.24
Volatile.Read + Volatile.ReadBarrier0.929 ns1.01
Volatile.Read + Interlocked.MemoryBarrier0.930 ns1.02

The barrier version costs about the same as the broken one. The read-everything-volatile version is about 24% slower, mostly because four ordered ldapur loads cannot be fused into two wide loads the way the plain copy can. Scale that up to a larger struct and the gap grows with the field count, while the barrier stays one instruction.

Two honest caveats. This is an uncontended, single-threaded loop: a dmb is cheap when the core has no outstanding memory traffic to wait for, which is why the full fence also looks free here. Under real write contention a full fence typically costs more than a load-only one, but I did not run a contended benchmark, so I am not putting a number on that. And all of this is arm64. On x64 both barriers emit no instruction, so you are only comparing what the JIT is allowed to do around them.

Gotchas that bite

Placement is everything. ReadBarrier orders reads before it against accesses after it. Putting it at the top of a reader, where people instinctively put a “volatile” read, orders nothing you care about. In a seqlock it goes after the data copy and before the second version read.

It is not a full fence. A store followed by ReadBarrier followed by a load can still be reordered. Dekker-style code, where each thread writes its own flag and then reads the other’s, needs Interlocked.MemoryBarrier() or an Interlocked operation.

It does not make anything atomic. The memory model spec is blunt: volatile semantics does not imply atomicity. If you skip the version validation, a barrier will happily order a torn read.

It is not a lock. A seqlock as written supports exactly one writer. Two writers need to serialize with an Interlocked.CompareExchange on the version (which is what GenericCache does) or with a real lock. If you are reaching for barriers because a lock felt slow, measure first: the post on lock vs Monitor vs SemaphoreSlim vs System.Threading.Lock shows how cheap an uncontended lock already is.

C# volatile fields are not the same tool. A volatile field makes the C# compiler emit every access with the volatile. IL prefix, so every read is an acquire and every write is a release. That is per-access Volatile.Read/Volatile.Write semantics, never a barrier over a batch, and it disables the load pairing shown above.

The verdict

Default to Volatile.Read. It is the right tool for nearly every lock-free flag, publication, and lazy-init pattern, it costs nothing on x64, and on modern arm64 it is a single load-acquire instruction. Use Volatile.ReadBarrier (on .NET 10 and later) only when the thing you need to order is a batch of earlier reads, typically a non-atomic copy that you validate afterwards. When you do, pair it with Volatile.WriteBarrier on the writer side, test on arm64, and remember that on arm64 WriteBarrier is as expensive as a full fence.

Sources

Comments

Sign in with GitHub to comment. Reactions and replies thread back to the comments repo.

< Back