A .PDX is the ADPCM sample bank a Sharp X68000 tune plays its percussion and voice hits from. An .MDX score carries no samples at all: it names a bank, and every hit lives in that separate file. The format is a pointer table and nothing else -- no magic, no version, no count, no names -- so like MDX, identifying one is arithmetic.
Every other part of the format follows from this. A slot is eight bytes at a fixed row, and the row index is the sample's identity: MML says a number and that is the row it reads. A bank whose only sample sits at row 33 carries ninety-five empty rows in front of it, and those rows are not waste to be tidied away -- removing one renumbers everything after it and silently retunes every tune that plays the bank.
Big-endian offset and length, eight bytes per row, ninety-six rows. An unused row is eight zero bytes. Sample data begins immediately after the last row.
Ninety-six rows of eight bytes is one bank, 768 bytes. Offsets are absolute file positions, not relative to anything -- the one place this format is simpler than the MDX that references it.
Nothing in the file states how many banks there are. A bank holding more than ninety-six samples simply repeats the table, and the sample data starts after the last one, so the count is recovered the same way MDX recovers its channel count:
An offset that is not a whole number of banks past zero describes no table at all, which is what makes the arithmetic an identification rather than a guess.
Rows may hold the same offset and the same length. That is a bank mapping one hit to several numbers so a part can play it without changing sample, and it is the normal case rather than a defect.
What does not happen is a row pointing part-way into another row's sample. Sharing is whole or not at all; a bank aliases, it does not slice. Anything reading a bank has to count distinct regions, not filled rows, or a shared hit is claimed twice and the file measures larger than it is.
The sample data is OKI MSM6258V ADPCM, which is the chip the X68000 has. Four bits per sample, one nibble per step, so a row's length in bytes is twice the number of audio samples it addresses.
The delta is the datasheet's, not a shortcut. Some decoders compute
the delta as ((2·d + 1) · step) >> 3, which is the same thing
up to rounding. The predictor integrates, so the rounding is not the same thing: an X68000
encoder modelled the chip, and only the per-term form brings a recorded hit back to silence
at its end. The other drifts by hundreds of units per sample.
There is no rate, no loop point, no root key and no name. A bank is an addressable pile of nibbles, and everything about how a sample is meant to sound lives in the MDX that plays it.
A filled row. Both values are unsigned and big-endian, and the offset is measured from the start of the file.
There is no magic number and no header before the table, so the table's own shape is the only claim the file makes about itself. Recognising one means reading ninety-six rows, discarding the empty ones, checking every remaining row lands inside the file, and checking the smallest offset is a whole number of banks.
Banks were also distributed packed with the X68000 compressors of
the day, and unlike a packed MDX nothing survives: a PDX is all table, so a compressed one
has no readable text to fall back on. The compressor stamps its name a few bytes in
-- LZX, ZOO, LHA or LZS --
which names what was done to the file but proves nothing about what is underneath it.
Both words in every row are big-endian, following the 68000 the machine is built around. The sample data itself is nibbles and has no byte order to get wrong.