Apple's answer to the 4 GB problem, and the third answer in a family that already had two. CAF keeps RIFF's four-character chunk ids and its payload-only sizes, but widens them to signed 64-bit, writes every field big-endian, and abandons the alignment rule entirely. Where RIFF pads chunks to 2 and Wave64 to 8, CAF chunks simply abut. A reader carrying either habit walks into the next chunk's id.
The size is signed, and -1 is legal. That is the whole reason the field is
an s64 rather than a u64. A writer streaming to a pipe does not know how
long the audio will be, and CAF lets it say so: a data chunk whose size is
-1 runs to the end of the file. Read as unsigned, that same chunk claims
18,446,744,073,709,551,615 bytes, which is the failure mode the signedness exists to
prevent and the one a reader ported from RIFF hits first.
The container is big-endian; the samples are not necessarily. Every
structural field -- the version, every chunk id, every size -- is big-endian throughout. The
sample data is whatever desc says it is, in a flags word that is independent of
the container's own byte order. A reader that infers one from the other is right only by luck, and wrong
silently, because wrong-endian PCM decodes to noise rather than to an error.
There is no alignment rule. A chunk ends where its payload ends and the next id begins on the very next byte. No pad, no rounding, no exceptions -- so a three-byte payload is followed immediately by a chunk id at an odd offset.
Eight bytes, and only the first four carry information a reader keys on. Unlike
RIFF there is no size here and no form type: the file's structure begins immediately with the
first chunk, and desc is required to be that chunk.
Thirty-two fixed bytes, required first, and the only chunk whose absence stops a
reader dead. Two of its fields are the ones a port from another format gets wrong: the sample rate
is a double rather than an integer, and format_flags decides the sample byte
order independently of the container.
The data chunk's payload does not begin with samples. Its first four
bytes are an edit count, and the audio starts after them. Everything downstream -- duration,
frame count, a carve of the raw stream -- is wrong by four bytes if that is missed, which for
16-bit mono is a two-sample shift and for a carve is four bytes of metadata glued to the front of
the audio.
| id | what it carries | notes |
|---|---|---|
desc | rate, codec, channels, bit depth, packet geometry | required, and required first |
data | the audio, after a u32 edit count | size -1 means "to the end of the file" |
pakt | packet count, valid frames, priming and remainder | required when bytes_per_packet is 0 |
chan | channel layout tag, bitmap, per-channel descriptions | which speaker, not how many |
info | a u32 count then NUL-terminated key/value strings | free-form metadata |
peak | per-channel peak amplitude and the frame it lands on | a float and a u64 per channel |
free | reserved space | room to grow the metadata without rewriting |
kuki | codec magic cookie | opaque decoder configuration |
| RIFF / WAVE | Wave64 | CAF | |
|---|---|---|---|
| chunk id | 4 chars | 16-byte GUID | 4 chars |
| size field | u32 | u64 | s64, signed |
| size counts | payload | payload + its 24-byte header | payload |
| alignment | 2 bytes | 8 bytes | none |
| endian | little | little | big |
| sample endian | little | little | a flag in desc |
| unknown length | -- | -- | size = -1 |
Three of those rows are places a reader written for one member produces a confident wrong answer on another rather than an error, which is the argument for reading the size field's type as carefully as its value.