What the file says about itself

A document still opened decades from now is found again by what it says about itself, not by the folder it happened to be in.

Summary

A PDF describes itself in a description packet that never appears on the page: the title, who prepared it, what it is about, when it dates from, and which rules it follows. Archives, search engines and filing software all read that packet.

Which rules it follows is written for you: the engine takes it from the standard the document claims. The rest you state yourself. A document with no title is a document an archive files under its file name, which is how a folder ends up holding four hundred files called the same thing.

Technically

The description is written as an XMP metadata packet, and the engine writes that packet whenever the document claims a standard. Written by the engine: the part, the level, the year of the edition, and the XMP extension schemas a checker needs before it accepts property names the standard does not already know — an electronic invoice, an accessibility claim and a print claim each add their own, and so does an XMP schema of your own.

Written from what you state, and omitted when you state nothing: the title, the author, the subject, the producer, the program the original document was written in, the words the document is to be found by, the date — which is written both as the creation date and as the modification date — and whether the document was trapped, that is whether overlaps have been added between adjoining inks so a slight misregistration on press leaves no white line. No archival rule requires any of these entries, so a document that leaves one out is still written. A document that claims the accessibility standard without a title is refused before a byte is written.

Parts 1 to 3 also carry the older document information dictionary, derived from the same values. Part 4 carries none: if the calling program sets an entry in the information dictionary under part 4, the engine refuses to write the document. Under part 3, an entry you write yourself in the information dictionary replaces the one derived from the same values, and does not flow back into the packet: where you fill in both places, keep the same value in each.

Clause by clause

Three things are not written: an identifier for the document, an identifier for this particular copy of it, and the moment the description itself was last touched. Where an archive requires any of them, an XMP schema of your own is the way in, and its property names are declared in the packet so a checker accepts them.

A packet handed over whole — for a page, an image, a drawing or an imported part — is written as it stands: nothing is added to it and nothing about it is checked. The document's own packet is always written from the facts you state.

What the engine does on its own

The engine does all of this as soon as a document claims the standard, whatever else the calling program asks for.

  • Writes the description packet for every document that claims a standard, with the part, the level and the year of the edition.
  • Declares, inside the packet, every property name the standard does not already know, so a checker accepts it.
  • Derives the older document information dictionary from the same values under parts 1 to 3, and writes none at all under part 4.

What you supply

The engine states what it was given and invents nothing else: without those values it refuses the document rather than choosing one for you.

  • The title, which is the one entry worth insisting on.
  • The author, the subject and the producer.
  • The date, in the ISO 8601 written form, stated by the calling program rather than read off the machine clock.
  • Any XMP schema of your own an archive requires, with a description for each of its properties.

The published texts behind these rules

Each published text, and what it asks for
Published text What it asks for
ISO 19005-4 6.1.3 Under part 4, a document says what it is in its metadata packet and nowhere else: it carries no information dictionary at all.

The other families of requirements

Which part of the archiving standard to claim Every font the document uses is embedded in the file Colours that mean the same thing on every device A page that comes out the same for everyone The document says what its marks stand for A document that carries other files inside it Checking the claim, before and after

Back to what the archiving standard asks for

See the prices See the examples