Talk to an engineer
INDUSTRIES/MEDIA & NEWS

Archives that answer, tooling built with the newsroom

Media organisations sit on decades of material they cannot search and deadlines that never move. We build retrieval and workflow tooling that respects rights, standards and the fact that editorial judgement is not automatable.

Book a technical callSee where we start

What the teams we work with are dealing with

ARCHIVE

Material nobody can find

Years of footage, copy and images searchable only by filename and the memory of whoever filed it.

WORKFLOW

Production overhead

Tagging, versioning, transcripts and rights checks eating hours that should go to reporting.

RIGHTS

Uncertain provenance

Content whose licence terms live in a contract nobody reads before publishing.

OUR WORK HERE

Where we usually start

Semantic archive search

Retrieval across text, transcripts and metadata so a researcher finds the clip by describing it, with the source shown.

Metadata and tagging

Automatic enrichment of incoming material, reviewed in bulk rather than item by item.

Editorial tooling

Interfaces built with the desk that uses them, for drafting support, versioning and fact-check trails.

Rights-aware generation

Anything generated is bounded by what your licences permit and labelled so an editor knows what they are looking at.

ENGAGEMENT

Typical shape of a media engagement

Short cycles, visible progress, and a scope you can change. You see working software every week rather than a status report.

WHAT YOU KEEP

The repository, the infrastructure code, the evaluation set and the documentation. In your accounts, under your licence, from the first commit.

First week

Time with the desk and the archive team, a sample corpus, and the questions people actually try to answer.

Second week

Working retrieval over a slice of the archive, measured against searches your researchers ran last month.

From there

Enrichment pipelines, editorial interfaces and rights checks, rolled out to one desk first.

Then on

Corpus expanded, quality reviewed with editorial, and tooling extended as the workflow changes.

What we are careful about

Editors decide

No automated publishing. Everything the system produces is a draft with a named human accountable for it.

Attribution

Outputs carry their sources, because in this industry an unsourced sentence is a liability.

Your standards

Tone, terminology and style guidance encoded from your own guide, not a generic one.

Licensing boundaries

What may be trained on, generated from or republished is agreed in writing before we build.

Tell us what the archive should be able to answer

Thirty minutes with an engineer who has shipped in this sector. You leave with a scope and a straight answer on feasibility.

Book a technical call