What the file says about itself

A document kept for thirty years is found again by what it says about itself, not by the folder it happened to be in.

In plain words

Written for anybody, with nothing taken for granted. It says what the rule is for and what it changes on the document you end up with.

On the spine of a binder somebody writes what is inside. A digital document does the same, in a block of text nobody ever sees on the page: the title, who prepared it, what it is about, when it dates from, and which rules it follows. Archives, search engines and filing software all read that block.

Which rules it follows is written for you — it comes straight from the claim. The rest is yours, because nobody else knows it. A document with no title is a document an archive files under its file name, which is how a folder ends up holding four hundred files called the same thing.

For a developer

The same thing without the detour: the objects, the constraints, and where the line falls between what is produced and what is declared.

The description is written as a metadata packet, and the packet is written whenever the document claims a standard. Written by the engine: the part, the level, the year of the edition, and the declarations a validator needs before it will accept property names the standard never heard of — an electronic invoice, an accessibility claim and a print claim each add their own, and so does a vocabulary of the caller's.

Written from what you state, and omitted when you state nothing: the title, the author, the subject, the producer, the date — which is written both as the creation date and as the modification date — and whether the job was trapped. No archival rule makes any of them compulsory, so none of them refuses the document; the accessibility standard asks for a title, and reports its absence.

Parts 1 to 3 also carry the older information dictionary, derived from the same values. Part 4 carries none, and an entry set there by hand refuses the document. An entry set by hand under part 3 overrides the derived one and does not flow back into the packet, so setting both means keeping both equal.

Clause by clause

The published text behind each requirement, what the engine measures the document against, and what it leaves to a validator.

Five things a reader might expect are not written, and should not be promised: an identifier for the document, an identifier for this particular copy of it, the moment the description itself was last touched, the name of the tool that composed the page, and keywords. Where an archive requires any of them, a vocabulary of your own is the way in, and its property names are declared in the packet so a validator accepts them.

The packet itself is written rather than inspected. A packet handed over whole by the caller is written as it stands, and nothing here measures it against the claim.

What the engine does on its own

Nothing on this list has to be asked for. It comes out of a document that claims the standard, whatever else the calling program says.

  • Writes the description packet for every document that claims a standard, with the part, the level and the year of the edition.
  • Declares, inside the packet, every property name the standard does not already know, so a validator accepts it.
  • Derives the older information block from the same values under part 3, and writes none at all under part 4.

What you supply

Nothing on this list can be guessed. The engine states what it was given and declines to invent the rest, which is why the document is refused rather than written with a value nobody chose.

  • The title, which is the one field worth insisting on.
  • The author, the subject and the producer.
  • The date, in the international written form, stated rather than read off the clock.
  • Any vocabulary of your own an archive requires, with a description for each of its properties.

The published texts behind it

Each published text, and what it asks for
Published text What it asks for
ISO 19005-4 6.1.3 A document of that part says what it is in its description and there alone.

The other families of requirements

Which part of the archiving standard to claim Every letter the document draws travels inside it Colours that mean the same thing on every device A page that comes out the same for everyone The document says what its marks stand for A document that carries other files inside it Checking the claim, before and after

Back to what the archiving standard asks for

See the prices See the examples

Glossary

standard
A rule argued out in committee, published under a number anybody may buy and read, and identical for every firm claiming it. A claim to follow one can therefore be checked against the text. A way of working that merely spread because it worked is a habit of the trade: useful, widespread, and answerable to no text at all.
veraPDF
A free program that checks whether a file really is what it claims to be, against the published rules. It is the tool archives use, so it is the one used here: a document is not called compliant because we say so, but because veraPDF passed it.
Factur-X
An invoice that is a page for a person and a data file for a machine, in one document. The page looks like any invoice; inside, the same amounts are attached in a form accounting software reads without anybody retyping them. French law requires this exchange between companies.
the identity card of a file
The block inside a file that says what the file is: its title, who made it, when, and which rules it follows. Search engines and archives read it; a reader never sees it. Its technical name is XMP.

Every word the site explains