This interactive analysis page was produced with Claude coding assistant. All content has been checked by a human reviewer. Errors or inconsistencies can be reported to issa@kcl.ac.uk
What NLS shared, the two segmentation tasks, and what the first pipeline run produced
400
catalogue films in two collections
—
videos processed ( Grampian tapes + other films)
—
programmes proposed by the pipeline
—
segments proposed inside those programmes
First pass. Everything the pipeline produced is represented here as it came out. We spot-checked a few items before and at the workshop, but besides these, the outputs have not been checked against the videos.
1 · The data NLS shared
Two collections, each catalogued in its own way
NLS shared 400 video files and their catalogue records. The two collections are described differently, so they raise different questions for segmentation.
Grampian — tapes
Grampian Television news tapes. One catalogue record per tape; its shotlist field lists items with start and end timecodes — the only item-level structure NLS holds for these tapes.
Non-Grampian — films
Films of many kinds (amateur, sponsored, educational, advertising, fiction…). One catalogue record per film; its shotlist is a prose, shot-by-shot description of a single film.
Hierarchical organisation
Mapping onto the design of FrameSense, we organised the collections hierarchically: archive → collection → video file → programme → segment. The bands below show, at each level, what the catalogue establishes on its own.
Band width = share of the 400 films, except “Working sample”: each square is one film from the hand-picked 16-film sample (solid = annotated by hand, hatched = not yet). Hover for numbers.
How much structured metadata is there?
Share of the 400 films with a value
Linked vocabularies (top four) versus item-level fields (bottom four)
Outside the shotlist there is almost no item-level structure: transmission date, episode title, producer and certificate are empty for practically every film. Genre and category terms cover under half the films, series terms very few.
2 · The two tasks
Outer segmentation finds the programmes in a video file; inner segmentation divides a programme into parts, sections, chapters, chunks of some kind...
ArchiveNLS
CollectionGrampian · non-Grampian
Video file processed
Programme proposed
Segment proposed
Outer segmentationGiven a video file, find each programme in it: where it starts and where it ends.
Inner segmentationGiven a programme, divide it into meaningful sections, each with its own timing and description.
What we have to compare against, per task
Grampian tapes
Non-Grampian films
Outer
Timed shotlist lines in the catalogue (consistent, but what a line stands for is unconfirmed). Hand-annotated programme boundaries for tapes.
The catalogue implies one programme per film, with a few “compilation” exceptions. Hand-annotated boundaries for films.
Inner
None yet: the timed lines describe items, not the parts inside a programme.
Prose shot-by-shot descriptions — usable, but not yet validated.
What output looks like
One example of each, drawn from the first-pass output: a typical Grampian tape (top), and a typical other film (bottom). Blocks are proposed programmes; the zoomed row shows one programme with its proposed segments. Only timings are shown — we have only checked against a handful of videos.
3 · What we processed
videos went through a vision-language pipeline
Input Grampian tapes and non-Grampian films, as delivered by NLS. (15 further non-Grampian films were not part of this run.)
→
PipelineKDL’s FrameSense pipeline drives an open vision-language model (Qwen3.8-27B) that watches each video and answers questions about it. Run on KCL’s research computing cluster.
→
OutputOne record per proposed programme, with its timing, descriptive fields and a segment list.
Model card — Qwen3.8-27B
Developer
Qwen (Alibaba)
Type
Dense vision-language model, 27 billion parameters; accepts text, images and video
Build used
4-bit quantised build (W4A16 GPTQ) published by RedHatAI, from Qwen/Qwen3.8-27B
Licence
Apache 2.0
Runs on
KCL’s research computing cluster (not a commercial API)
These were the field targeted in the first pass, as agreed on the prioritisation according to usefulness for NLS:
start & durationtitleyearplacecolourtype(s)fiction / non-fictionsummarysynopsissegment list (timecodes + labels)creditskeywords
The output covers programme boundaries (outer) and a list of parts inside each programme (partially inner), plus descriptive metadata for each programme.
4 · Results: outer segmentation
Tape files tend to be cut into many programmes (titles), non-Grampian files almost always map to a single programme (title) — and almost every video is accounted for from start to end
Programmes per Grampian file
Median , mean , largest
Programmes per non-Grampian file
of films came out as a single programme
Programme duration — Grampian
Median min · longer than 30 min (not shown) · under 10 s
Programme duration — non-Grampian
Median min · longer than 30 min (not shown) · under 10 s
How much of each video the programmes account for
Median ; lowest ; files below 95%
Gap between consecutive programmes
Median s · smallest s · overlaps · gaps over 30 s (not shown)
The pattern matches what the catalogue led us to expect, the cuts never overlap, and they are separated by a few seconds — consistent with the pipeline finding the breaks between programmes. It doesn't tell us if the cuts fall in the right places..
5 · Results: descriptive metadata and inner segments
Descriptions are nearly complete; titles and years mostly are not; labels do not yet match the catalogue
Fields left empty
Share of programmes with no value
Segments per programme
Share of programmes · median (Grampian), (non-Grampian)
Most frequent types
A programme can carry several types
Most frequent places
Places split into separate names; a programme can carry several
of Grampian programmes are labelled “tv news”
of Grampian programmes are labelled “documentary”
of Grampian tapes have no programme labelled “tv news”
6 · Data quality
The output needs tidying before it can be used in downstream tasks.
Data cleaning notes
Issue
What we found
What we did
Segment timecodes
Three styles, two of which look identical (HH:MM:SS and MM:SS:00). Read naively, programmes get segment times of hours inside programmes of under an hour.
Resolved per programme by checking against the programme’s own duration; a few programmes mix both styles and stay approximate.
Places
Free text with mixed delimiters and order: distinct strings, once order is ignored, distinct names — town, region and country mixed in one string.
Split into names. A proper place field needs an authority list.
Sparse fields
Year and title are empty for most programmes (see section 5).
Left empty; whether to generate them at all is a discussion point.
Unreadable items
segment entries had no readable timecode.
Kept, flagged.
File encoding
Windows-1252, not UTF-8.
Read accordingly.
7 · Findings
What we now know, what we could know with more development, and what needs a human first
What we now know
The pipeline ran end to end on videos and produced programmes with segments and descriptions.
of files trip none of our checks — which doesn't mean they are correct but are at least consistent.
What we could know with more development
Which errors are systematic (by collection, tape length, era) — and whether a re-run gives the same cuts.
Consistent labels: mapping types to NLS’s vocabulary and places to an authority list (e.g. a gazetteer).
What needs human checks before more development
Whether the cuts are right on the flagged files: too many, too few, fragments, hour-long “programmes”, long gaps.
Whether the type labels can be trusted (news tapes labelled documentary).
Whether summaries, places and years are factually right.
A few unflagged files as a control.
8 · Suggested human review plan
Sixteen kinds of oddity, ordered make the most of human watching
We turned the checks into simple rules and ranked the files that trip them. of files raise at least one flag ( high priority, medium); programmes () and gaps are flagged individually.
What is flagged, by rule
Darker = look first. Units differ by rule: programmes, gaps or files.
How the order is chosen: Each next file is the one that adds the most problem types not yet seen (a rule’s first few example files count in full, repeats only a little), so a reviewer meets every kind of problem early instead of the same one twenty times. The top 20 mix Grampian and non-Grampian files. Note that a flag means “worth a look”, not necessarily “wrong”.
9 · For discussion at WS1
Topics we think are worth discussing together
What counts as a programme?
A news item? A section of a tape? Each part of a compilation?
What is a timed line in a Grampian shotlist, and how was this field created?
Could those lines seed or check programme boundaries — or do they mean something else?
What should be produced at each level?
Which fields do we want per programme, and which per segment?
How do we judge quality?
How many seconds of error is acceptable at a boundary? Is there a reference for this or can we can agree on a benchmark?
What to do with non-programme material?
Dead air, bars, test signals, very short fragments: keep, label or drop?
Shared vocabulary
NLS’s terms versus the model’s (for example “tv sport” and “tv sports”); places against an authority list; keywords controlled or free?
How will corrections flow back?
What should a review interface let staff do — adjust in and out points, split or merge programmes, link segments to rights information?