ISSA workshop 1 — NLS

From catalogue to first-pass segmentation

This interactive analysis page was produced with Claude coding assistant. All content has been checked by a human reviewer. Errors or inconsistencies can be reported to issa@kcl.ac.uk

Data to drive this page is generated by the predictions_analysis.ipynb notebook.

What NLS shared, the two segmentation tasks, and what the first pipeline run produced

400
catalogue films in two collections
—
videos processed ( Grampian tapes + other films)
—
programmes proposed by the pipeline
—
segments proposed inside those programmes
1 · The data NLS shared

Two collections, each catalogued in its own way

NLS shared 400 video files and their catalogue records. The two collections are described differently, so they raise different questions for segmentation.

Grampian — tapes

Grampian Television news tapes. One catalogue record per tape; its shotlist field lists items with start and end timecodes — the only item-level structure NLS holds for these tapes.

Non-Grampian — films

Films of many kinds (amateur, sponsored, educational, advertising, fiction…). One catalogue record per film; its shotlist is a prose, shot-by-shot description of a single film.

Hierarchical organisation

Mapping onto the design of FrameSense, we organised the collections hierarchically: archive → collection → video file → programme → segment. The bands below show, at each level, what the catalogue establishes on its own.

Grampian — verifiedGrampian — estimated / assumed non-Grampian — verifiednon-Grampian — estimated / assumed ⚠ Needs human checkNot yet determined

Band width = share of the 400 films, except “Working sample”: each square is one film from the hand-picked 16-film sample (solid = annotated by hand, hatched = not yet). Hover for numbers.

How much structured metadata is there?

Share of the 400 films with a value

Linked vocabularies (top four) versus item-level fields (bottom four)

Outside the shotlist there is almost no item-level structure: transmission date, episode title, producer and certificate are empty for practically every film. Genre and category terms cover under half the films, series terms very few.

2 · The two tasks

Outer segmentation finds the programmes in a video file; inner segmentation divides a programme into parts, sections, chapters, chunks of some kind...

ArchiveNLS
CollectionGrampian · non-Grampian
Video file processed
Programme proposed
Segment proposed
Outer segmentationGiven a video file, find each programme in it: where it starts and where it ends.
Inner segmentationGiven a programme, divide it into meaningful sections, each with its own timing and description.

What we have to compare against, per task

Grampian tapesNon-Grampian films
OuterTimed shotlist lines in the catalogue (consistent, but what a line stands for is unconfirmed). Hand-annotated programme boundaries for tapes. The catalogue implies one programme per film, with a few “compilation” exceptions. Hand-annotated boundaries for films.
InnerNone yet: the timed lines describe items, not the parts inside a programme. Prose shot-by-shot descriptions — usable, but not yet validated.

What output looks like

One example of each, drawn from the first-pass output: a typical Grampian tape (top), and a typical other film (bottom). Blocks are proposed programmes; the zoomed row shows one programme with its proposed segments. Only timings are shown — we have only checked against a handful of videos.

3 · What we processed

videos went through a vision-language pipeline

Input Grampian tapes and non-Grampian films, as delivered by NLS. (15 further non-Grampian films were not part of this run.)
→
PipelineKDL’s FrameSense pipeline drives an open vision-language model (Qwen3.8-27B) that watches each video and answers questions about it. Run on KCL’s research computing cluster.
→
OutputOne record per proposed programme, with its timing, descriptive fields and a segment list.

Model card — Qwen3.8-27B

DeveloperQwen (Alibaba)
TypeDense vision-language model, 27 billion parameters; accepts text, images and video
Build used4-bit quantised build (W4A16 GPTQ) published by RedHatAI, from Qwen/Qwen3.8-27B
LicenceApache 2.0
Runs onKCL’s research computing cluster (not a commercial API)
SettingsVideo sampling, prompts and inference environment: see more technical details in our notes on video sampling

Model pages: Qwen/Qwen3.8-27B · RedHatAI/Qwen3.8-27B-INT4 (as read on 24 Sept 2026).

What comes out, per programme

These were the field targeted in the first pass, as agreed on the prioritisation according to usefulness for NLS:

start & durationtitleyearplacecolourtype(s)fiction / non-fictionsummarysynopsissegment list (timecodes + labels)creditskeywords

The output covers programme boundaries (outer) and a list of parts inside each programme (partially inner), plus descriptive metadata for each programme.

4 · Results: outer segmentation

Tape files tend to be cut into many programmes (titles), non-Grampian files almost always map to a single programme (title) — and almost every video is accounted for from start to end

Programmes per Grampian file

Median , mean , largest

Programmes per non-Grampian file

of films came out as a single programme

Programme duration — Grampian

Median min · longer than 30 min (not shown) · under 10 s

Programme duration — non-Grampian

Median min · longer than 30 min (not shown) · under 10 s

How much of each video the programmes account for

Median ; lowest ; files below 95%

Gap between consecutive programmes

Median s · smallest s · overlaps · gaps over 30 s (not shown)
The pattern matches what the catalogue led us to expect, the cuts never overlap, and they are separated by a few seconds — consistent with the pipeline finding the breaks between programmes. It doesn't tell us if the cuts fall in the right places..
5 · Results: descriptive metadata and inner segments

Descriptions are nearly complete; titles and years mostly are not; labels do not yet match the catalogue

Fields left empty

Share of programmes with no value

Segments per programme

Share of programmes · median (Grampian), (non-Grampian)

Most frequent types

A programme can carry several types

Most frequent places

Places split into separate names; a programme can carry several
of Grampian programmes are labelled “tv news”
of Grampian programmes are labelled “documentary”
of Grampian tapes have no programme labelled “tv news”
6 · Data quality

The output needs tidying before it can be used in downstream tasks.

Data cleaning notes
IssueWhat we foundWhat we did
Segment timecodesThree styles, two of which look identical (HH:MM:SS and MM:SS:00). Read naively, programmes get segment times of hours inside programmes of under an hour.Resolved per programme by checking against the programme’s own duration; a few programmes mix both styles and stay approximate.
PlacesFree text with mixed delimiters and order: distinct strings, once order is ignored, distinct names — town, region and country mixed in one string.Split into names. A proper place field needs an authority list.
Sparse fieldsYear and title are empty for most programmes (see section 5).Left empty; whether to generate them at all is a discussion point.
Unreadable items segment entries had no readable timecode.Kept, flagged.
File encodingWindows-1252, not UTF-8.Read accordingly.
7 · Findings

What we now know, what we could know with more development, and what needs a human first

What we now know

  • The pipeline ran end to end on videos and produced programmes with segments and descriptions.
  • of files trip none of our checks — which doesn't mean they are correct but are at least consistent.

What we could know with more development

  • Which errors are systematic (by collection, tape length, era) — and whether a re-run gives the same cuts.
  • Consistent labels: mapping types to NLS’s vocabulary and places to an authority list (e.g. a gazetteer).

What needs human checks before more development

  • Whether the cuts are right on the flagged files: too many, too few, fragments, hour-long “programmes”, long gaps.
  • Whether the type labels can be trusted (news tapes labelled documentary).
  • Whether summaries, places and years are factually right.
  • A few unflagged files as a control.
8 · Suggested human review plan

Sixteen kinds of oddity, ordered make the most of human watching

We turned the checks into simple rules and ranked the files that trip them. of files raise at least one flag ( high priority, medium); programmes () and gaps are flagged individually.

What is flagged, by rule

Darker = look first. Units differ by rule: programmes, gaps or files.
How the order is chosen: Each next file is the one that adds the most problem types not yet seen (a rule’s first few example files count in full, repeats only a little), so a reviewer meets every kind of problem early instead of the same one twenty times. The top 20 mix Grampian and non-Grampian files. Note that a flag means “worth a look”, not necessarily “wrong”.
9 · For discussion at WS1

Topics we think are worth discussing together

What counts as a programme?

A news item? A section of a tape? Each part of a compilation?

What is a timed line in a Grampian shotlist, and how was this field created?

Could those lines seed or check programme boundaries — or do they mean something else?

What should be produced at each level?

Which fields do we want per programme, and which per segment?

How do we judge quality?

How many seconds of error is acceptable at a boundary? Is there a reference for this or can we can agree on a benchmark?

What to do with non-programme material?

Dead air, bars, test signals, very short fragments: keep, label or drop?

Shared vocabulary

NLS’s terms versus the model’s (for example “tv sport” and “tv sports”); places against an authority list; keywords controlled or free?

How will corrections flow back?

What should a review interface let staff do — adjust in and out points, split or merge programmes, link segments to rights information?