One bit per sample, at sixty-four times the CD rate. Direct Stream Digital
is what a Super Audio CD holds, and it does not store amplitudes at all: each sample is a single
bit from a sigma-delta modulator, and the signal lives in the local density of ones. Two
containers carry it, one from each company behind the disc, and they agree about almost
nothing.
the one fact that breaks every other reader
bits per sample is 1. Every duration and size formula written for PCM
assumes at least eight, and the assumption is usually invisible because it is spelled
bits / 8 somewhere in a helper. The sample count is a count of bits per
channel. The byte count is that divided by eight, times the channel count. Reading the two as
the same number is not a small error; it is a factor of eight, and it produces a duration that
looks entirely plausible.
At 2,822,400 Hz -- the DSD64 rate, sixty-four times 44,100 -- one
second of stereo is 5,644,800 bits, which is 705,600 bytes. A reader that assumes 16-bit
samples expects 11,289,600 for that same second and finds a sixteenth of it. The duration
survives, because seconds are still samples over rate; it is every calculation that touches
bytes that breaks -- the buffer size, the seek target, the offset of the next block.
DSF: the DSD chunk (28 bytes)
Sony's container. Flat, little-endian, four blocks in fixed order and no nesting
at all. The header says how long the file is and where its tag lives.
DSF: the fmt block (first 32 of 52 bytes)
Everything about the audio, and two fields that look like the same number and are
not: channel type is a layout id, channel num is a count.
▸channel type is not a channel count
Type 4 is quad -- front left, front right, back left, back right. Type 5 is also four channels, and they are front left, front right, centre and LFE. The count is the same and the speakers are not, so a reader that derives one field from the other silently puts the centre channel in the back of the room. The two are separate fields precisely because they are separate facts, and a file whose type and count disagree is a file to distrust.
The full table: 1 mono, 2 stereo, 3 three channels, 4 quad, 5 four channels, 6 five channels, 7 5.1.
▸the last block is padded, not short
Block size per channel is fixed at 4,096 bytes and the spec is explicit about the tail: "Block size per channel is fixed, so please fill ZERO(0x00) for unused sample data area in the block." The final block is therefore full-length and zero-filled, not truncated. A reader that treats the whole final block as audio emits a burst of silence rather than noise, which is the merciful failure -- but a reader that computes the length from the block count rather than from the sample count overstates the duration.
DSDIFF: FRM8 and the version chunk (28 bytes)
Philips' container, and an IFF file in every respect but one: the size is
64-bit. The spec says so plainly -- "this FORM chunk is slightly different from EA IFF 85
(the ckDataSize is not a long but a double ulong)". Big-endian throughout, and the even-length pad
byte is kept.
▸FRM8 counts from byte 12, and Wave64 does not
Three formats widened IFF to survive past 4 GB and each drew the line somewhere different. DSDIFF keeps the IFF convention: the size counts everything after the id and the size itself, so the file is size + 12. Wave64 counts its own 24-byte header inside the size, so the chunk is exactly size bytes. RF64 keeps 32-bit sizes and moves the real length into a separate ds64 chunk, leaving 0xFFFFFFFF as a sentinel behind.
All three are reasonable and no two are the same, which is why a reader written for one of them is wrong about the others in a way that still produces a number.
DSDIFF: PROP and its first local chunk (28 bytes)
DSDIFF nests. PROP is a container whose payload opens with a property type
-- SND for sound -- followed by local chunks that describe the audio. The
sample rate, the channel list, the compression type and the absolute start time all live in
here rather than at the top level.
the local chunks inside PROP
FSsample rate, u32. 2,822,400 is DSD64; each higher rate doubles
CHNLa u16 count then one 4-character id per channel: SLFT SRGT for stereo, MLFT MRGT C LFE LS RS for multichannel
CMPRa 4-character compression id, a length byte, then a human-readable name. DSD is uncompressed; DST is the lossless codec
ABSSabsolute start time: hours u16, minutes u8, seconds u8, samples u32. Where this file sits in the original recording
LSCOloudspeaker configuration, u16. 0 stereo, 3 five channels, 4 five-point-one
▸DST, and why a DSDIFF can be lossless-compressed
A Super Audio CD holds 4.7 GB and a stereo DSD64 stream runs at 5.6 Mbit/s, so an uncompressed disc would hold well under two hours and a multichannel one far less. Direct Stream Transfer is the lossless codec that closes the gap, and a DSDIFF carrying it replaces the DSD sound chunk with DST , holding DSTF frames, an optional DSTC CRC per frame and a DSTI index.
The compression type in CMPR is what says which of the two a file is, and it is the only thing that does: both look identical at the container level.
where the tag lives
Both containers embed ID3v2, the tag format MP3 made ubiquitous, and they
put it in completely different places. DSF stores a 64-bit pointer in its header to
a tag at the end of the file, and sets it to zero when there is none. DSDIFF carries an
ID3 chunk, and in practice writes it inside PROP alongside the sample
rate rather than at the top level.
This is the same tag a RIFF file carries in an id3 chunk and an
AIFF in an ID3 chunk. Five containers, one tag format, and five different
opinions about where it goes.
the rate family
DSD642,822,400 Hz -- 64 x 44,100. The SACD rate
DSD1285,644,800 Hz
DSD25611,289,600 Hz
DSD51222,579,200 Hz
48 kHz base3,072,000 and its doublings. Legal, rarer, and the reason a rate should be named from a table rather than divided by 44,100