ACIDCAT . FILE FORMAT REFERENCE

CAF Anatomy

Apple Core Audio FormatMac OS X 10.4 . 2005
rev 2026.09
magic caff
ids 4-char
sizes s64, signed
endian big

Apple's answer to the 4 GB problem, and the third answer in a family that already had two. CAF keeps RIFF's four-character chunk ids and its payload-only sizes, but widens them to signed 64-bit, writes every field big-endian, and abandons the alignment rule entirely. Where RIFF pads chunks to 2 and Wave64 to 8, CAF chunks simply abut. A reader carrying either habit walks into the next chunk's id.

the three places it differs

The size is signed, and -1 is legal. That is the whole reason the field is an s64 rather than a u64. A writer streaming to a pipe does not know how long the audio will be, and CAF lets it say so: a data chunk whose size is -1 runs to the end of the file. Read as unsigned, that same chunk claims 18,446,744,073,709,551,615 bytes, which is the failure mode the signedness exists to prevent and the one a reader ported from RIFF hits first.

The container is big-endian; the samples are not necessarily. Every structural field -- the version, every chunk id, every size -- is big-endian throughout. The sample data is whatever desc says it is, in a flags word that is independent of the container's own byte order. A reader that infers one from the other is right only by luck, and wrong silently, because wrong-endian PCM decodes to noise rather than to an error.

There is no alignment rule. A chunk ends where its payload ends and the next id begins on the very next byte. No pad, no rounding, no exceptions -- so a three-byte payload is followed immediately by a chunk id at an odd offset.

the header

Eight bytes, and only the first four carry information a reader keys on. Unlike RIFF there is no size here and no form type: the file's structure begins immediately with the first chunk, and desc is required to be that chunk.

the audio description

Thirty-two fixed bytes, required first, and the only chunk whose absence stops a reader dead. Two of its fields are the ones a port from another format gets wrong: the sample rate is a double rather than an integer, and format_flags decides the sample byte order independently of the container.

the audio

The data chunk's payload does not begin with samples. Its first four bytes are an edit count, and the audio starts after them. Everything downstream -- duration, frame count, a carve of the raw stream -- is wrong by four bytes if that is missed, which for 16-bit mono is a two-sample shift and for a carve is four bytes of metadata glued to the front of the audio.

the chunks
idwhat it carriesnotes
descrate, codec, channels, bit depth, packet geometryrequired, and required first
datathe audio, after a u32 edit countsize -1 means "to the end of the file"
paktpacket count, valid frames, priming and remainderrequired when bytes_per_packet is 0
chanchannel layout tag, bitmap, per-channel descriptionswhich speaker, not how many
infoa u32 count then NUL-terminated key/value stringsfree-form metadata
peakper-channel peak amplitude and the frame it lands ona float and a u64 per channel
freereserved spaceroom to grow the metadata without rewriting
kukicodec magic cookieopaque decoder configuration
where the family diverges
RIFF / WAVEWave64CAF
chunk id4 chars16-byte GUID4 chars
size fieldu32u64s64, signed
size countspayloadpayload + its 24-byte headerpayload
alignment2 bytes8 bytesnone
endianlittlelittlebig
sample endianlittlelittlea flag in desc
unknown length----size = -1

Three of those rows are places a reader written for one member produces a confident wrong answer on another rather than an error, which is the argument for reading the size field's type as carefully as its value.