Skip to content

Archive indexing

Digital records index review across a scanned records archive

A digital indexing review measures whether the index over a scanned records archive actually retrieves the documents it claims to hold. It is run for digital records managers during a migration, after a vendor handoff, or when searches keep missing documents the team knows exist. Sampled index rows are traced to image files, OCR text, and metadata fields, and every row that fails to find, open, or correctly describe its document is logged. The result is an exception list the team can use to correct the index rather than rescan the archive.

When this review is needed

  • A scanning project has delivered images and a spreadsheet index, and nobody has tested one against the other.
  • Two archives with different indexing conventions are being merged into one system.
  • Searches for known documents return nothing, or return another aircraft's paperwork.
  • An upcoming transaction will expose the archive to an external review team for the first time.

The problem

An archive is only as findable as its index. Scanning projects are priced per page, so indexing gets done quickly, from cover sheets and file names, often by staff with no records background, and the defects are invisible at handover because the images all exist. The records team inherits a system that demonstrates well and fails quietly, one unretrievable document at a time.

What gets reviewed

  • Index field population: which rows carry document type, date, aircraft, and reference fields and which sit empty
  • Row-to-file linkage across the sampled index
  • OCR text quality for the fields that search depends on
  • Aircraft and engine attribution in metadata against sampled document content
  • Naming and foldering conventions across scanning batches and vendors
  • Coverage: documents present as images but absent from the index

Scope this review

Tell us the asset, the event, and the evidence in scope, and we will outline a focused first engagement.

Send a representative, redacted record set and we will scope the review.

What gets validated

  • A stratified sample of index rows opens the correct file on the first attempt
  • Document dates and types in the index match what the sampled image actually shows
  • Registration and serial numbers in metadata agree with the document face
  • Full-text search retrieves sampled documents by the identifiers a records auditor would use
  • Un-indexed files found in the folder tree are counted and characterized

Evidence normally required

  • The current index export, with field definitions if they exist
  • Read access to the image repository with original file paths
  • OCR output files or the search layer built on them
  • Vendor delivery notes or batch manifests from the scanning project
  • A short list of documents the team already knows are hard to find

Common discrepancies

  • Index rows pointing at renamed or moved files that now open nothing
  • Metadata copied down a spreadsheet column, giving hundreds of documents the same date
  • OCR that reads stamps and handwriting as noise, leaving key records invisible to search
  • Whole scanning batches present on disk but never merged into the index

What is at stake

Index defects surface exactly when retrieval matters: an auditor's sample, a lease-return demand, an AD accomplishment question. Each miss forces a manual hunt through folder trees, erodes trust in the archive, and pushes people back toward keeping paper, which defeats the migration the index was supposed to complete.

Move from findings to resolution

Move from findings to a documented resolution path.

How the work runs

01

Profile the index

Characterize field population, conventions, and batch boundaries across the full index export before drawing any sample.

02

Draw a stratified sample

Select rows by scanning batch, document type, and date range, weighted toward strata where defects are likely to cluster.

03

Trace rows to files and content

Open each sampled row's target, verify the document matches the metadata, and test retrieval through the search layer.

04

Rank the corrections

Report defect rates per stratum and order corrections by their effect on retrieval, so effort lands where searches actually fail.

What the buyer receives

  • A row-level exception list stating what failed for each sampled index entry
  • A defect profile by scanning batch and vendor, showing where problems cluster
  • Recommended index corrections ranked by retrieval impact

Who uses the output

  • Digital records managers correcting the index and setting acceptance rules for future batches
  • Records control leads deciding which parts of the archive can be relied on today
  • Teams preparing the archive for a transaction data room

How the work fits into the transaction or program

Indexing review usually runs first in an archive quality program, because every content-level review depends on retrieval: a redelivery binder or AD accomplishment record that search cannot reach might as well not exist. Its defect profile also sets acceptance criteria for future scanning batches, so the archive stops accumulating the same faults.

Jurisdiction-specific considerations

FAA guidance on electronic recordkeeping, including AC 120-78, expects electronic systems to preserve the integrity and retrievability of the records they replace, and 14 CFR 91.417 still governs what must be kept, with AC 43-9C describing what adequate entries look like. EASA operators face equivalent retention duties under Regulation (EU) 1321/2014. The review checks retrievability and attribution against those expectations without asserting compliance on the operator's behalf.

Regulatory limits

The review evaluates the index, the metadata, and search behavior over sampled records. It makes no airworthiness determination, does not certify the archive or the recordkeeping system, and does not approve electronic records for use in place of paper.

What this review does not cover

  • Building or rebuilding the index itself
  • Selecting or configuring document-management software
  • Content review of the underlying maintenance records beyond attribution sampling

Specific to this review

  • A scan that search cannot reach behaves exactly like a lost document, even though the image sits safely on disk.
  • Index defects cluster by batch: one bad vendor week can quietly corrupt thousands of rows while the rest of the archive stays clean.
  • OCR confidence scores, where the vendor supplied them, let sampling target the rows most likely to be wrong instead of sampling blind.
  • Merging two archives multiplies defects, because each system's assumptions about naming and dates survive the merge.

Sources

Frequently asked questions

Do you review every row in the index?

No. The review samples, stratified by scanning batch, document type, and date range, because defects cluster by how the archive was built. The sample is sized to state defect rates per stratum with useful confidence, which is what a correction plan needs.

Relevant glossary terms

Related pages

Where this fits

Talk to an engineer who has done this work

We will walk through your current state, the records or evidence involved, and a scoped first engagement.

Talk through the aircraft, records, evidence, deadline, and next useful step.