Skip to content
NiyathnairPublic

About

Does the video reach the words? A frozen language model finishes a sentence about a clip it is watching, and controls show the video is what does it. 1.3 M trained parameters, tested on Physion against people.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

CLIF-GPT

Does the video reach the words?
A frozen language model finishes a sentence about a clip it is watching,
and a set of controls shows that the video is what does it.

tests MIT licence 1.3 M trained parameters distilgpt2 and SigLIP are frozen

One sentence, two clips. For the clip where the red object reaches the yellow mat the model continues with touched or hit; for the clip where it does not, with avoided or missed.

One sentence, two clips the model never trained on. The language model is frozen and never saw a caption.
On its own it continues with "the". Watching the left clip it writes touched. Watching the right one, avoided.
This pair was picked as a clear example. Every number below is over all 1,200 clips.

What this is

CLIF-GPT couples a frozen language model, distilgpt2, with a small trained video branch. The branch watches a clip and adds one vector to the single hidden state from which the language model reads its next word. Nothing else is touched: the language model, its vocabulary and the vision encoder all stay frozen. The trained part has 1.3 million parameters, against 82 million in the language model.

The question this repository answers is whether that is enough for the video to reach the words. It is easy to build a demo that looks as if it does. It is harder to show it. So the sentences here are written to give nothing away, and every result is paired with controls: no video, a blank video, and another clip's video.

It does reach the words. Watching the whole clip, the frozen language model picks the right outcome word on 83 % of clips it never trained on, and the right word is the single most likely token out of 50,257. With any control it is at 50 %. The effect carries to sentences and words that were never in training. Two things I hoped for did not hold, and they are reported with the same weight: from the first 1.5 seconds alone the model is ten points behind people, and a learned "imagined future" did not help.

Results at a glance

All on Physion's 1,200 test clips with five-fold cross-validation, so every clip is scored by a model that never saw it. Three seeds; ± is the standard deviation across them.

Question Measured
Does the video reach the words? 83.2 ± 0.6 % right when the model watches the clip. 50.0 % for the language model alone or with a blank video, 51.4 % with another clip's video.
Is the right word merely preferred, or is it the top word? It is the top word of the whole vocabulary on 83.1 % of clips. On the same sentences the language model alone never puts an outcome word first.
Does speaking through a frozen language model cost accuracy? No. The same video branch with a plain yes/no head gets 82.8 %.
Does it carry to sentences it never trained on? Yes, fully: 83.0 % on new wordings, 82.9 % on everyday sentences about something else.
Does it carry to words it never trained on? The direction does, for all seven untrained word pairs (AUC 0.885 to 0.891, against 0.891 for trained pairs). The size of the push does not always beat the language model's own preference: 50 to 77 % at a fixed threshold.
Does it carry to a kind of scene it never trained on? Reporting what it sees, mostly: 76.2 % with the scenario held out of training. Predicting, no: 56.8 %.
Can it predict, from the 1.5 seconds people were shown? 64.0 ± 0.5 %. People get 74.2 %.
Does imagining the future help it predict? No: 63.6 % with a learned rollout of the future, against 64.0 % without.

The test

The eight Physion scenarios, each shown at the start, after 1.5 seconds and at the end.

Physion (Bear et al., NeurIPS 2021) is 1,200 short simulated clips in eight scenarios. In each clip one object is painted red and one object or zone is painted yellow. The label is one bit: did red touch yellow before everything came to rest? People were shown the first 1.5 seconds and asked to predict it; they were right 74.2 % of the time.

Each clip is paired with a sentence that stops one word short of the outcome:

The red object moved toward the yellow one and ___

The next word depends on the clip: hit or touched if there was contact, missed or avoided if not. The sentence is the same either way, so a model that ignores the video scores exactly 50 %. That is the property the first version of this repository lacked.

Four sets of language were fixed before any training (cilf/prompts.py):

Set Example Used for
8 training sentences, 2 word pairs "The red thing headed for the yellow area and" → hit / missed training, on other clips
8 new wordings "By the end of the video the red object had" never trained on
4 new word pairs struck / passed, reached / dodged, bumped / cleared, met / skipped never used as targets
8 everyday sentences, 3 word pairs "The cyclist rode straight at the gate and" → crashed / stopped never trained on

Two protocols. In watch, the model sees the whole clip, outcome included: the question is whether what the video shows gets into the sentence. In predict, it sees the first 1.5 seconds, as people did: the question is whether it can anticipate.

How it works

flowchart LR
  subgraph FROZEN[" frozen "]
    direction TB
    V["SigLIP vision encoder"]
    L["distilgpt2"]
    W["its output matrix"]
  end
  C["clip"] --> V --> T["tokens per moment:<br/>scene, red object, yellow object,<br/>their relation, a close-up"]
  T --> E["temporal encoder<br/>2 layers"]
  S["sentence"] --> L --> H["hidden states"]
  E --> X["cross-attention<br/>video and sentence"]
  H --> X
  X --> B["bias b<br/>rank 64"]
  H --> A(("h + b"))
  B --> A --> W --> N["next word"]
  classDef frozen fill:#eef1f4,stroke:#8c959f,color:#24292f
  classDef trained fill:#ddebfb,stroke:#2a78d6,color:#0b2e5c
  class V,L,W frozen
  class T,E,X,B trained
Loading

Only the blue boxes are trained. The bias starts at exactly zero, so an untrained model is the language model. Training asks for one thing: by cross-entropy over the whole vocabulary, make the next word the right outcome word.

One frame in its two renders, the masks their difference gives, and the close-up crop.

Where the objects come from. Physion ships every clip twice, in ordinary colours and with the two objects painted red and yellow. The two renders differ only on those objects, so their difference is an exact mask of each in every frame. No detector is used and no colour threshold is tuned. Each moment of a clip then becomes five tokens: the frame's embedding, the vision features under each mask, a close-up of the two objects, and their geometry, including five cues about whether red stands on yellow or merely in front of it.

Everything the vision encoder produces is computed once and cached, so training one model takes under two minutes on a laptop.

1. Does the video reach the words?

Dot charts. Whole clip: language model alone 50 %, blank video 50 %, another clip's video 51 %, the clip's own video 83 %. First 1.5 seconds: 50, 50, 50.5 and 64 %, with people at 74 %.
whole clip first 1.5 s
language model alone 50.0 50.0
+ a blank video 50.0 ± 0.0 50.0 ± 0.0
+ another clip's video 51.4 ± 0.8 50.5 ± 0.4
+ the clip's own video 83.2 ± 0.6 64.0 ± 0.5
right word is the top word of 50,257 83.1 ± 0.6 64.1 ± 0.5
references without a language model
same video branch, yes/no head 82.8 ± 0.1 63.6 ± 0.6
linear probe on the frozen features 80.5 60.6
threshold on the smallest on-screen gap 71.3 52.3
people 74.2

Three readings.

The video is doing it. Give the model another clip's video and it falls to chance. Give it a blank one and it says the same word every time.

The frozen language model costs nothing. The same branch with an ordinary yes/no head, trained the same way, gets 82.8 %. Routing the answer through a frozen 82-million-parameter model and its 50,257-word vocabulary loses no accuracy.

Why 83 % and not 100 %, when the outcome is on screen? Because overlapping on screen is not touching. In 59 % of the no-contact clips the red and yellow masks overlap at some point, since one is simply in front of the other, and a rule that thresholds the on-screen gap gets 71 %. Telling the two apart from frozen features and 960 training clips is where the remaining errors are. The limit is perception, not language: the yes/no head stops at the same place.

2. Can it predict?

Per scenario, accuracy from the first 1.5 seconds: the model against people. People are ahead in seven of eight scenarios.

Four clips drawn at random. The model watches 1.5 seconds, writes the next word, and the clip then plays on. Three are right and one is wrong.

Four held-out clips drawn at random with a fixed seed, mistakes included. Three right, one wrong.

From the first 1.5 seconds the model is right on 64.0 % of clips. People are right on 74.2 %. It is ahead of people in one scenario of eight, cloth draping, and well behind on rolling and dropping. It also does not find the same clips hard that people do: across clips, the correlation between how often it is right and how often people are is 0.18.

Imagining the future did not help. The first version of this repository was built around a "video ODE" that would simulate what happens next. Here that idea gets a fair test: a latent ODE rolls the scene forward from the last observed moment, it is trained to match the frozen features of the real future, and the language model reads the imagined future along with the observed past. Accuracy is 63.6 %, against 64.0 % without it. The rollout learns to continue the features but adds nothing the first 1.5 seconds did not already give.

This is where the claim stops. A small branch on frozen features can report physics it sees. It does not yet anticipate physics as well as a person does.

3. Does it transfer?

Area under the ROC curve: the chance that a contact clip gets higher log-odds for the contact word than a no-contact clip. Whole clip, three seeds. In brackets, accuracy with the threshold fixed at zero.

sentences ↓ / outcome words → trained on never targets story words
trained on 0.891 (83 %) 0.887 (62 %) 0.888 (61 %)
new wordings 0.891 (83 %) 0.886 (62 %) 0.889 (58 %)
everyday stories 0.891 (83 %) 0.887 (62 %) 0.889 (63 %)

New sentences cost nothing. "The cyclist rode straight at the gate and" was never in training and has nothing to do with the clip. Shown a clip with contact, the model continues it with hit; shown one without, with missed, as reliably as for the sentences it trained on.

New words move the right way, for a reason that can be written down. The video adds b to a hidden state, so the log-odds of any two words u, v shift by exactly (w_u − w_v) · b, where w are rows of the language model's own output matrix (docs/bound.md). The video therefore pushes a direction, and every pair of words whose difference leans that way moves with it. All seven untrained pairs do:

word pair AUC accuracy at zero cosine to the trained direction
hit / missed (trained) 0.891 83 % +0.78
touched / avoided (trained) 0.890 83 % +0.78
reached / dodged 0.888 77 % +0.29
slammed / escaped 0.891 73 % +0.25
met / skipped 0.888 59 % +0.24
struck / passed 0.885 52 % +0.18
bumped / cleared 0.886 60 % +0.09
crashed / stopped 0.889 59 % +0.08
smashed / halted 0.886 50 % +0.07

The ranking is as good for untrained words as for trained ones. The accuracy is not, and the table shows why: the further a pair's direction is from the trained one, the smaller the push, and for smashed / halted it never overturns the language model's own preference. The video moves meaning; it was only trained to win for four words.

A kind of scene never seen in training. Train on seven scenarios, test on the eighth:

scenario whole clip, scenario in training whole clip, scenario held out first 1.5 s, in training first 1.5 s, held out
collide 89.0 89.8 ± 1.6 68.8 57.8 ± 5.1
contain 80.4 71.7 ± 2.0 64.0 57.8 ± 2.2
dominoes 96.7 76.5 ± 5.9 60.4 54.1 ± 1.8
drape 71.3 59.4 ± 2.3 71.1 54.5 ± 4.2
drop 80.2 79.7 ± 0.9 58.8 54.3 ± 3.9
link 81.3 74.9 ± 4.0 56.2 54.3 ± 1.0
roll 72.9 70.5 ± 1.0 66.0 62.0 ± 2.9
support 93.5 86.9 ± 1.2 66.8 59.5 ± 3.5
all 83.2 76.2 ± 0.5 64.0 56.8 ± 2.1

Watching the whole clip, most of the skill carries over: 76.2 % on a kind of scene the model never trained on, against 83.2 % when it did. Collisions, drops and rolling lose almost nothing. Dominoes lose the most, 97 % to 77 %, and cloth draping ends lowest, at 59 %. A plausible reason is that nothing in the other seven scenarios looks like a chain of falling pieces or like cloth settling on an object. Predicting from 1.5 seconds does not carry over: 56.8 %, a few points above chance. Whatever the model uses to anticipate an outcome is specific to the scenarios it trained on.

4. What the video branch needs

Accuracy of the full model and of variants with one part removed or replaced.
variant accuracy change AUC on untrained words
full model 83.2 ± 0.6 0.886
no close-up of the two objects 81.7 ± 0.6 -1.5 0.874
no object geometry 80.7 ± 0.4 -2.5 0.869
whole-frame embedding only 68.6 ± 3.0 -14.5 0.756
object geometry only 78.0 ± 0.1 -5.2 0.854
no time order: the clip as a bag of moments 83.2 ± 0.5 +0.1 0.894
bias from the video alone, no sentence 82.8 ± 0.4 -0.3 0.892
min-plus fusion, from the first version 84.2 ± 0.3 +1.0 0.866

Whole clip, three seeds each. What each row says:

  • It needs to look at the objects. With only the embedding of the whole frame, accuracy falls from 83 % to 69 %. The frozen vision encoder's summary of a frame does not say whether two small things in it are touching.
  • Geometry carries most of it. The boxes of the two objects and the contact cues alone reach 78 %. Vision features on top of them add five points; the close-up adds a point and a half of those.
  • The order of moments does not matter. Take away the time embeddings and nothing changes. To say whether a touch happened, a bag of moments is enough; knowing which came first adds nothing. This was only tested with the whole clip in view.
  • Neither does reading the sentence. A bias computed from the video alone does as well as one that cross-attends to the sentence. Every sentence here asks the same thing, so there is nothing for that attention to pick up. Whether it helps when sentences ask different things is untested.
  • Min-plus fusion is not worse. I expected the first version's min(h, b + w) to lose. It is a point ahead on the trained words, 84.2 against 83.2, and behind on untrained ones, AUC 0.866 against 0.886. Addition stays the default because it was fixed before this run, and because it is the fusion whose effect on any word can be written down.

How far can the video move the language model?

The fusion is additive, which gives two exact statements (docs/bound.md, with proofs):

  • The log-odds of any two words shift by (w_u − w_v) · b. A zero bias leaves the model untouched.
  • The whole next-word distribution moves by at most min(Δ, Δ²/8) nats of KL divergence, where Δ is the spread of the score change over the vocabulary.

Measured over all 86,400 clip and sentence pairs of the main run: the video moves the distribution by 10.6 nats on average, the bound averages 24.7, and it is never violated.

What changed from the first version

The first version of this repository described the same idea and did not test it. In order of importance:

  1. The answer never depended on the video. Every prompt was a template with a fixed target word, so a lookup on the prompt alone scored 100 % on the training, validation and transfer files. Here the sentence never determines the answer.
  2. The transfer test rewarded ignoring the video. Each "held-out" prompt was paired with a random clip from every scenario and expected the same word each time. One of them was in the training set verbatim.
  3. Labels were per scenario. Physion's own per-clip outcome was in the data folder, unused. It is now the label.
  4. No control. The model was compared with the bare language model only. There is now a shuffled-video and a blank-video control on every result.
  5. The demo's answer was partly a lookup. It matched the prompt against training captions with hard-coded keyword rules and could replace the model's word with a clip's caption label. The demo here prints the model's own top words and nothing else.
  6. The theorem was circular, and the script it cited was missing. docs/bound.md replaces it with a statement that is proved and tested.
  7. "Sheaf alignment" and "tropical fusion" were the norm of a difference of two linear maps and an element-wise minimum. Neither was explained or measured. The first is gone. The second is now measured, as one row of the ablation, and it holds up: a point ahead on trained words, behind on untrained ones.

Also removed: the slot-attention and YOLO paths, the webcam scripts and the per-scenario caption manifests. The first version is kept at the tag v0.3-prototype.

Quick start

Requirements: Python 3.10 or newer, ffmpeg on the path, and about 3 GB of disk for the data and cached features. Tested on macOS with Apple Silicon; the tests also run on Linux.

git clone https://github.com/Niyathnair/clif-gpt.git
cd clif-gpt
python -m pip install -r requirements.txt
python -m pytest tests -q                 # 28 tests; one waits for distilgpt2 to be cached and is skipped until then

Try a trained model on a clip it never saw:

python scripts/download_physion.py        # 284 MB
python scripts/demo.py --list 8           # held-out clips of the checkpoint
python scripts/demo.py --clip pilot_it2_collision_assorted_targets_box_0001
clip      pilot_it2_collision_assorted_targets_box_0001  (collide; truth: the red object touched the yellow one)
model saw the whole clip
sentence  The red object moved toward the yellow one and ___

language model alone   the 23%   then 7%   it 4%   moved 3%   left 2%   was 2%
with the video         touched 54%   hit 46%   avoided 0%   struck 0%   missed 0%   moved 0%

Pass --sentence "..." for any sentence of your own, and --checkpoint checkpoints/predict.pt for the model that sees only the first 1.5 seconds.

Reproduce the numbers

python scripts/download_physion.py        # the clips
python scripts/extract_features.py        # frozen vision features, about an hour on a laptop GPU, 1.8 GB
bash experiments/run_all.sh               # every run in this README, about two hours
python experiments/pair_detail.py         # the per-word-pair table
python experiments/make_tables.py         # results/tables.md
python experiments/make_figures.py        # assets/fig_*.png
python scripts/make_demo_media.py         # the two demo clips

One run on its own, for example the main one:

python experiments/run.py watch --view watch

results/*.json holds the metrics of every run reported here. Hyperparameters were chosen once, on an inner split of the first fold's training clips, and then left alone; five settings were tried and all landed within a point of each other.

Repository layout

cilf/
  physion.py      the clips, their labels, human accuracy, the cross-validation folds
  video.py        frames, object masks from the two renders, geometry and contact cues
  features.py     frozen vision features, computed once and cached
  store.py        cached features as fixed-size tensors: watch, prefix, future
  prompts.py      the sentences and outcome words, fixed in advance
  lm.py           the frozen language model: cached sentence states and its output matrix
  model.py        the trained video branch, the fusion, the latent ODE
  train.py        training and scoring
  metrics.py      forced choice, AUC, top word
  baselines.py    gap rule, linear probe, yes/no classifier
  bound.py        the KL bound
experiments/      run.py (one variant, cross-validated), run_all.sh, tables, figures
scripts/          download, feature extraction, the demo, the demo media
tests/            28 tests
docs/bound.md     the identity and the bound, with proofs
results/          metrics of every reported run
checkpoints/      one trained model per protocol, each with the list of clips it never saw

Limits, and what is not new

  • Simulated clips, one question. Physion is rendered, the objects are colour-coded, and the only thing asked is whether two of them touch. Nothing here is evidence about real video or open-ended physical reasoning.
  • The masks come from the benchmark. They need Physion's two renders. On other video this would need a tracker.
  • Small data. Each model trains on 960 clips. The clips are Physion's test set, used with cross-validation because that is where the per-clip labels and human data are. This is not Physion's official protocol and these numbers are not leaderboard entries.
  • Below people at prediction, and no better with a learned rollout.
  • Four trained words. The video's push is strong enough to win for the words it trained on and not always for others.
  • Small language models only. Swapping distilgpt2 for gpt2-medium, four times the size, changes nothing: 83.1 % watching and 64.2 % predicting. That fits the limit being perception, and says nothing about a modern language model.

Not new: adding a learned vector to a frozen language model's hidden state, cross-attention between video and text, neural ODEs for latent dynamics, and the Physion benchmark are all existing work. What this repository adds is a clean test of whether the video reaches the words, with the controls that the question needs, and the per-word account of why the effect transfers.

Credits

  • Physion: D. Bear et al., Physion: Evaluating Physical Prediction from Vision in Humans and Machines, NeurIPS 2021 Datasets and Benchmarks. The clips, the labels and the human accuracy per clip are theirs, including the label file in data/physion/. The clips are downloaded from the authors' bucket and are not redistributed here; the short excerpts on this page are shown to illustrate results.
  • SigLIP: X. Zhai et al., Sigmoid Loss for Language Image Pre-Training, ICCV 2023.
  • distilgpt2: Hugging Face's distilled GPT-2.

Licence

The code is released under the MIT licence.

About

Does the video reach the words? A frozen language model finishes a sentence about a clip it is watching, and controls show the video is what does it. 1.3 M trained parameters, tested on Physion against people.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages