adding vision to SOTA text-only coding models
TL;DR
- Laguna-S-2.1 is a strong code and reasoning model that cannot see. It gained sight through a trained 35.4M parameter projector, with nothing else changed.
- The vision tower is Qwen3-VL's, frozen. The language model is Laguna's, frozen and byte-identical to the released weights. Compute cost was about $95.
- On MMMU-Pro it scores 41.0%. Replace its image features with noise and it drops to 13.67%, below the 25% chance floor for four-option questions.
- Feeding it garbage makes it worse than guessing, which is the strongest evidence available that the image features carry real information.
35.4M
trained parameters, the only ones
41.0%
MMMU-Pro, best checkpoint
13.67%
same eval, image swapped for noise
~$95
4x B200, 3.3 hours
the method
A vision tower turns pixels into a sequence of embeddings. A language model turns token ids into embeddings of its own. Both are learned representations of meaning, and neither knows about the other. Learning the map between them lets a model that has never seen an image read one.
So freeze both and train only the map. The projector is 35.4M parameters.
norm LayerNorm(1152)
2x2 spatial concat -> 4608
linear_fc1 Linear(4608, 4608)
GELU
linear_fc2 Linear(4608, 3072)1152 is the tower's hidden size, 3072 is Laguna's. The 2x2 concat merges four adjacent image patches into one token, which cuts sequence length by four and costs little, since adjacent patches are highly redundant.
Image tokens enter through placeholder ids in the prompt. At the embedding step, the projector's output overwrites the embeddings at those positions. Everything downstream is stock Laguna.
The tower's own merger provided a partial head start. Its first two layers map from the same 1152-dim space, so those weights transfer directly. The final layer maps to 4096 and Laguna's hidden size is 3072, so it starts random. Four of six tensors transferred.
training data
66,000 short-answer visual question pairs, drawn from 21 sources across natural images, scene text, charts, documents, diagrams, counting and tables. Captions and long descriptions are excluded by construction, not filtered out afterwards.
Long targets are off-policy for a frozen language model. They carry high loss no matter how well the vision features are aligned, which drowns out the signal of interest. Short answers keep the loss sensitive to whether the image is being used.
loss cannot tell you whether it works
The targets are short answers like Answer: B. A model can drive loss down by learning that format and ignoring the image completely. The loss curve looks the same either way, and it looks healthy.
So every evaluation runs twice. Once normally, once with the image features replaced by Gaussian noise of the same shape and scale, with everything else held fixed. Three outcomes are possible.
- blind equals sighted. The model ignores images.
- blind near chance. The model uses images but can fall back on question text.
- blind below chance. The image features are load-bearing, and noise actively misleads.
Run the same evaluation early and late in training and the control does the work.
| samples | sighted | blind | gap |
|---|---|---|---|
| 6,400 | 32.50% | 35.00% | -2.50 |
| 57,600 | 41.00% | 13.67% | 27.33 |
At 6,400 samples the model is answering from question wording. Four-option questions often have eliminable distractors, which is why it beats 25% without looking at anything. By 57,600 samples the same questions produce a 27 point gap, about ten standard errors from zero at n=300.
Nothing in the loss curve distinguishes those two states cleanly. The control does it in one number.
two epochs was worse than one
The run trained 2070 steps, two full passes over the data. Training loss more than halved. Held-out accuracy fell.
| checkpoint | sighted | blind | gap | train loss | samples |
|---|---|---|---|---|---|
| step 900 | 41.00% | 13.67% | 27.33 | 1.2957 | 57,600 |
| step 2070 | 38.33% | 13.00% | 25.33 | 0.6149 | 132,480 |
Both deltas sit near the standard error of 2.7, so the claim is not that the second epoch actively hurt. The claim is it clearly did not help and cost twice as much. The steep part of that loss drop began exactly when the run crossed into its second pass over the same 66,000 samples, which is what memorisation looks like.
Judged on loss alone the worse checkpoint would have shipped.
benchmarks
Every benchmark runs twice. Once normally, once with the image features replaced by Gaussian noise of the same shape and scale. The gap is the measurement. A sighted score on its own cannot distinguish a model reading images from one answering off the question wording.
sighted vs blind, n=300 per benchmark
MMMU-Progap 25.33 · chance 25%
MMMU_DEV_VALgap 25.33 · chance 25%
MMBench_DEV_ENgap 48.00 · chance 25%
HallusionBenchgap 47.34 · chance 50%
TextVQA_VALgap 71.33 · chance n/a
OCRVQAgap 37.67 · chance n/a
SEEDBench_IMGgap 49.00 · chance 25%
The four-option sets have a 25% floor. HallusionBench is yes/no, so its floor is 50%, which means a sighted score near 51% is at chance and not a capability result.
Where a blind score lands below chance on a four-option set, noise in the image slots is actively misleading the model rather than being ignored. HallusionBench behaves differently. Its blind score of 4.33% is far below the 50% floor, which almost certainly means the output becomes unparseable under noise rather than confidently inverted. Its gap is therefore inflated relative to the multiple-choice gaps and should not be compared directly against them.
TextVQA and OCRVQA are open-ended and scored by normalised containment against the reference answers, which is looser than official VQA accuracy. Their absolute numbers are not comparable to published leaderboards. The gap remains meaningful because both arms use the identical scorer.
All numbers come from the shipped checkpoint, which measured 2.67 points below the best checkpoint recorded during training. That earlier checkpoint was lost to a retention policy before it could be kept.
cost
0.203 s/sample 4x B200 effective batch 64 900 steps 3.3 hours about $95
Cheap enough that the interesting constraint stops being money and starts being how quickly you can tell whether a run is working. That is an argument for evaluating during training rather than at the end. Evaluations ran at 6,400, 57,600 and 132,480 samples, and the middle one turned out to be the checkpoint worth keeping.
three things that broke
- Vision towers with near-identical architecture can have incompatible preprocessing. Two towers built on the same base were tested, both 1152 hidden, 27 layers, 4304 intermediate width, 16 heads. One uses patch 16 with a temporal dimension and emits flat 1536-wide patch vectors. The other uses patch 14 and emits shaped [n, 3, 14, 14] tensors. Feeding one tower the other's output fails immediately, and the similarity of the model configs makes it easy to assume otherwise.
- Generation breaks the naive splicing implementation. During prefill the prompt carries the image placeholder tokens, so features and slots match. Every decode step afterwards passes only the newest token under the KV cache. That token contains zero placeholders while the projector output is still held, so an assertion comparing features to slots fires on the second token. The fix is to permit the zero-slot decode case while keeping the assertion for genuine mismatches, since a genuine mismatch fails silently and produces a model reading misaligned features.
- Checkpoint retention does not know which checkpoint is good. Step 900 measured as the best result, then a keep-the-last-three rule deleted it while training continued. The shipped checkpoint is 2.67 points below the measured peak. Pin checkpoints that evaluate well, before the next save cycle.
what shipped
both are public.
weights
mm_projector (35.40M, the only trained part), frozen vision tower (576.39M), configs and tokenizer files.
huggingface.cocode
Training, evaluation with blind controls, and the 12-line projector definition.
github.comThe text backbone is not included, because it is byte-identical to the upstream release. That is also why the no-regression claim is structural rather than measured. With no image tokens in the prompt, the forward pass is bit-identical to stock Laguna-S-2.1. Every backbone weight is frozen, the projector is not in the text path, and the placeholder id never appears.
how to run
Three parts have to be assembled at load time, the stock language model, the frozen vision tower, and the trained projector. The projector is not a transformers architecture, so the module is defined by hand. It is 12 lines.
the projector
import torch, torch.nn as nn
class Projector(nn.Module):
def __init__(self, vis_hidden=1152, merge=2, text_hidden=3072, eps=1e-6):
super().__init__()
inner = vis_hidden * merge * merge # 4608
self.norm = nn.LayerNorm(vis_hidden, eps=eps)
self.linear_fc1 = nn.Linear(inner, inner)
self.act = nn.GELU()
self.linear_fc2 = nn.Linear(inner, text_hidden)
self.inner = inner
def forward(self, x): # x: [n_patches, 1152]
x = self.norm(x)
x = x.reshape(-1, self.inner) # 2x2 spatial merge
return self.linear_fc2(self.act(self.linear_fc1(x)))inference
from PIL import Image
@torch.no_grad()
def ask(image: Image.Image, question: str, max_new_tokens: int = 64) -> str:
feat = proc(images=image.convert("RGB"), return_tensors="pt")
grid = feat["image_grid_thw"]
n_img = int(grid.prod(-1).sum() // 4) # merge_size ** 2 == 4
q_ids = tok(question, add_special_tokens=False).input_ids
ids = torch.tensor([[IMAGE_TOKEN_ID] * n_img + q_ids], device=DEV)
feats = tower(feat["pixel_values"].to(DEV, torch.bfloat16),
grid_thw=grid.to(DEV))
if hasattr(feats, "last_hidden_state"):
feats = feats.last_hidden_state
img_emb = proj(feats) # [n_img, 3072]
emb = lm.get_input_embeddings()(ids)
mask = (ids == IMAGE_TOKEN_ID).unsqueeze(-1).expand_as(emb)
assert int(mask[..., 0].sum()) == img_emb.shape[0]
emb = emb.masked_scatter(
mask, img_emb.reshape(-1, emb.shape[-1]).to(emb.dtype))
out = lm.generate(inputs_embeds=emb, max_new_tokens=max_new_tokens,
do_sample=False, pad_token_id=tok.eos_token_id or 1)
return tok.decode(out[0], skip_special_tokens=True).strip()Full loading code, including the tower and processor setup, is in the repo.
obstacles to running
- Generate with
inputs_embeds, notinput_ids. The image is spliced into embeddings, so there is no token id that carries it. Passing input_ids to generate silently produces a text-only answer. - Use this release's preprocessor. The tower expects patch 16 with a temporal dimension, emitting flat 1536-wide patch vectors. A processor from a different tower will produce shaped tensors of the wrong width and fail immediately, or worse, silently mis-tile.
- Do not hardcode the placeholder count. It follows the image resolution. Compute it from
image_grid_thwevery time, and assert it against the mask before scattering. That assertion is the cheapest protection against silently misaligned features. - Keep prompts short and answers short. The model was trained on short-form visual question answering with captions excluded. A request for long-form description tests something the model never saw.
sanity check
Run any prompt twice, once normally and once with the image features replaced by noise of the same scale. If the two answers are equally good, the model is not reading the image and something in the wiring is wrong. It is the same control used to validate the model during training, and it takes one line.
img_emb_blind = torch.randn_like(img_emb) * img_emb.std()limitations
- Supervised finetuning on short answers only. No reinforcement learning stage is in these weights, and long-form description about images was never trained and is weak.
- Hallucination is the clearest weakness. HallusionBench sits at its 50% floor. The 47-point gap says the model reads the image. The absolute score says it does not reliably contradict a false premise in the question.
- Fine text is limited by the image token budget. Patches merge 2x2 before the projector, so dense documents lose resolution before the language model sees them.
next steps
reinforcement learning at a measurable scale
The model answers immediately rather than reasoning about what it sees. A GRPO run of 5,120 rollouts produced a reward curve with no detectable learning.
| metric | mean | std | slope / step | trend |
|---|---|---|---|---|
| Reward | 0.1185 | 0.0220 | -0.00041 ± 0.00322 | flat |
| Correct | 0.0879 | 0.0209 | -0.00003 ± 0.00306 | flat |
The run is too short to count as evidence against RL. A 64-prompt batch cannot resolve the effect, because the step-to-step noise is wider than any improvement a short run could produce. At 2.5 seconds per rollout, a run long enough to see a trend is roughly $500 to $900. Anyone repeating this should size the run against the noise, not against a step count.
unfreeze the language model
The largest structural limit is that the backbone has no circuits for reasoning over visual input. A LoRA or partial unfreeze is the standard second stage. Backward would then flow through 5.3B active parameters instead of 35.4M, so expect roughly 3 to 5 times the original training cost, still a few hundred dollars at these throughputs.
raise the image token budget for fine text
TextVQA (scene text, large and central) scored 74.00% sighted while OCRVQA (book and CD covers, small, stylised, rotated) scored 43.67%, at comparable training volume. The 30 point spread comes from resolution, not alignment. Images are capped at 1,024 tokens. Raising that cap, or using a higher-resolution tower configuration, is the direct test.
more alignment data
Training used 66,000 short-answer pairs. Standard alignment sets are several hundred thousand. A second epoch over the existing data made held-out accuracy worse, so the constraint is data variety rather than more passes over the same rows.