back

adding vision to SOTA text-only coding models

  • coding models
  • vision
  • finetuning

TL;DR

  • Laguna-S-2.1 is a strong code and reasoning model that cannot see. It gained sight through a trained 35.4M parameter projector, with nothing else changed.
  • The vision tower is Qwen3-VL's, frozen. The language model is Laguna's, frozen and byte-identical to the released weights. Compute cost was about $95.
  • On MMMU-Pro it scores 41.0%. Replace its image features with noise and it drops to 13.67%, below the 25% chance floor for four-option questions.
  • Feeding it garbage makes it worse than guessing, which is the strongest evidence available that the image features carry real information.

35.4M

trained parameters, the only ones

41.0%

MMMU-Pro, best checkpoint

13.67%

same eval, image swapped for noise

~$95

4x B200, 3.3 hours

the method

A vision tower turns pixels into a sequence of embeddings. A language model turns token ids into embeddings of its own. Both are learned representations of meaning, and neither knows about the other. Learning the map between them lets a model that has never seen an image read one.

So freeze both and train only the map. The projector is 35.4M parameters.

norm        LayerNorm(1152)
            2x2 spatial concat -> 4608
linear_fc1  Linear(4608, 4608)
GELU
linear_fc2  Linear(4608, 3072)

1152 is the tower's hidden size, 3072 is Laguna's. The 2x2 concat merges four adjacent image patches into one token, which cuts sequence length by four and costs little, since adjacent patches are highly redundant.

Image tokens enter through placeholder ids in the prompt. At the embedding step, the projector's output overwrites the embeddings at those positions. Everything downstream is stock Laguna.

The tower's own merger provided a partial head start. Its first two layers map from the same 1152-dim space, so those weights transfer directly. The final layer maps to 4096 and Laguna's hidden size is 3072, so it starts random. Four of six tensors transferred.

training data

66,000 short-answer visual question pairs, drawn from 21 sources across natural images, scene text, charts, documents, diagrams, counting and tables. Captions and long descriptions are excluded by construction, not filtered out afterwards.

Long targets are off-policy for a frozen language model. They carry high loss no matter how well the vision features are aligned, which drowns out the signal of interest. Short answers keep the loss sensitive to whether the image is being used.

loss cannot tell you whether it works

The targets are short answers like Answer: B. A model can drive loss down by learning that format and ignoring the image completely. The loss curve looks the same either way, and it looks healthy.

So every evaluation runs twice. Once normally, once with the image features replaced by Gaussian noise of the same shape and scale, with everything else held fixed. Three outcomes are possible.

Run the same evaluation early and late in training and the control does the work.

samplessightedblindgap
6,40032.50%35.00%-2.50
57,60041.00%13.67%27.33

At 6,400 samples the model is answering from question wording. Four-option questions often have eliminable distractors, which is why it beats 25% without looking at anything. By 57,600 samples the same questions produce a 27 point gap, about ten standard errors from zero at n=300.

Nothing in the loss curve distinguishes those two states cleanly. The control does it in one number.

two epochs was worse than one

The run trained 2070 steps, two full passes over the data. Training loss more than halved. Held-out accuracy fell.

checkpointsightedblindgaptrain losssamples
step 90041.00%13.67%27.331.295757,600
step 207038.33%13.00%25.330.6149132,480

Both deltas sit near the standard error of 2.7, so the claim is not that the second epoch actively hurt. The claim is it clearly did not help and cost twice as much. The steep part of that loss drop began exactly when the run crossed into its second pass over the same 66,000 samples, which is what memorisation looks like.

Judged on loss alone the worse checkpoint would have shipped.

benchmarks

Every benchmark runs twice. Once normally, once with the image features replaced by Gaussian noise of the same shape and scale. The gap is the measurement. A sighted score on its own cannot distinguish a model reading images from one answering off the question wording.

sighted vs blind, n=300 per benchmark

MMMU-Progap 25.33 · chance 25%

38.33%
13.00%

MMMU_DEV_VALgap 25.33 · chance 25%

45.33%
20.00%

MMBench_DEV_ENgap 48.00 · chance 25%

73.67%
25.67%

HallusionBenchgap 47.34 · chance 50%

51.67%
4.33%

TextVQA_VALgap 71.33 · chance n/a

74.00%
2.67%

OCRVQAgap 37.67 · chance n/a

43.67%
6.00%

SEEDBench_IMGgap 49.00 · chance 25%

68.00%
19.00%

The four-option sets have a 25% floor. HallusionBench is yes/no, so its floor is 50%, which means a sighted score near 51% is at chance and not a capability result.

Where a blind score lands below chance on a four-option set, noise in the image slots is actively misleading the model rather than being ignored. HallusionBench behaves differently. Its blind score of 4.33% is far below the 50% floor, which almost certainly means the output becomes unparseable under noise rather than confidently inverted. Its gap is therefore inflated relative to the multiple-choice gaps and should not be compared directly against them.

TextVQA and OCRVQA are open-ended and scored by normalised containment against the reference answers, which is looser than official VQA accuracy. Their absolute numbers are not comparable to published leaderboards. The gap remains meaningful because both arms use the identical scorer.

All numbers come from the shipped checkpoint, which measured 2.67 points below the best checkpoint recorded during training. That earlier checkpoint was lost to a retention policy before it could be kept.

cost

0.203 s/sample   4x B200    effective batch 64
900 steps        3.3 hours  about $95

Cheap enough that the interesting constraint stops being money and starts being how quickly you can tell whether a run is working. That is an argument for evaluating during training rather than at the end. Evaluations ran at 6,400, 57,600 and 132,480 samples, and the middle one turned out to be the checkpoint worth keeping.

three things that broke

what shipped

both are public.

The text backbone is not included, because it is byte-identical to the upstream release. That is also why the no-regression claim is structural rather than measured. With no image tokens in the prompt, the forward pass is bit-identical to stock Laguna-S-2.1. Every backbone weight is frozen, the projector is not in the text path, and the placeholder id never appears.

how to run

Three parts have to be assembled at load time, the stock language model, the frozen vision tower, and the trained projector. The projector is not a transformers architecture, so the module is defined by hand. It is 12 lines.

the projector

import torch, torch.nn as nn

class Projector(nn.Module):
    def __init__(self, vis_hidden=1152, merge=2, text_hidden=3072, eps=1e-6):
        super().__init__()
        inner = vis_hidden * merge * merge          # 4608
        self.norm       = nn.LayerNorm(vis_hidden, eps=eps)
        self.linear_fc1 = nn.Linear(inner, inner)
        self.act        = nn.GELU()
        self.linear_fc2 = nn.Linear(inner, text_hidden)
        self.inner      = inner

    def forward(self, x):                            # x: [n_patches, 1152]
        x = self.norm(x)
        x = x.reshape(-1, self.inner)                # 2x2 spatial merge
        return self.linear_fc2(self.act(self.linear_fc1(x)))

inference

from PIL import Image

@torch.no_grad()
def ask(image: Image.Image, question: str, max_new_tokens: int = 64) -> str:
    feat = proc(images=image.convert("RGB"), return_tensors="pt")
    grid = feat["image_grid_thw"]
    n_img = int(grid.prod(-1).sum() // 4)            # merge_size ** 2 == 4

    q_ids = tok(question, add_special_tokens=False).input_ids
    ids   = torch.tensor([[IMAGE_TOKEN_ID] * n_img + q_ids], device=DEV)

    feats = tower(feat["pixel_values"].to(DEV, torch.bfloat16),
                  grid_thw=grid.to(DEV))
    if hasattr(feats, "last_hidden_state"):
        feats = feats.last_hidden_state
    img_emb = proj(feats)                            # [n_img, 3072]

    emb  = lm.get_input_embeddings()(ids)
    mask = (ids == IMAGE_TOKEN_ID).unsqueeze(-1).expand_as(emb)
    assert int(mask[..., 0].sum()) == img_emb.shape[0]
    emb  = emb.masked_scatter(
        mask, img_emb.reshape(-1, emb.shape[-1]).to(emb.dtype))

    out = lm.generate(inputs_embeds=emb, max_new_tokens=max_new_tokens,
                      do_sample=False, pad_token_id=tok.eos_token_id or 1)
    return tok.decode(out[0], skip_special_tokens=True).strip()

Full loading code, including the tower and processor setup, is in the repo.

obstacles to running

sanity check

Run any prompt twice, once normally and once with the image features replaced by noise of the same scale. If the two answers are equally good, the model is not reading the image and something in the wiring is wrong. It is the same control used to validate the model during training, and it takes one line.

img_emb_blind = torch.randn_like(img_emb) * img_emb.std()

limitations

next steps

reinforcement learning at a measurable scale

The model answers immediately rather than reasoning about what it sees. A GRPO run of 5,120 rollouts produced a reward curve with no detectable learning.

metricmeanstdslope / steptrend
Reward0.11850.0220-0.00041 ± 0.00322flat
Correct0.08790.0209-0.00003 ± 0.00306flat

The run is too short to count as evidence against RL. A 64-prompt batch cannot resolve the effect, because the step-to-step noise is wider than any improvement a short run could produce. At 2.5 seconds per rollout, a run long enough to see a trend is roughly $500 to $900. Anyone repeating this should size the run against the noise, not against a step count.

unfreeze the language model

The largest structural limit is that the backbone has no circuits for reasoning over visual input. A LoRA or partial unfreeze is the standard second stage. Backward would then flow through 5.3B active parameters instead of 35.4M, so expect roughly 3 to 5 times the original training cost, still a few hundred dollars at these throughputs.

raise the image token budget for fine text

TextVQA (scene text, large and central) scored 74.00% sighted while OCRVQA (book and CD covers, small, stylised, rotated) scored 43.67%, at comparable training volume. The 30 point spread comes from resolution, not alignment. Images are capped at 1,024 tokens. Raising that cap, or using a higher-resolution tower configuration, is the direct test.

more alignment data

Training used 66,000 short-answer pairs. Standard alignment sets are several hundred thousand. A second epoch over the existing data made held-out accuracy worse, so the constraint is data variety rather than more passes over the same rows.