Weakly Aligned: Why VLMs Can't Learn Language From a Baby's-Eye View

1. Babies learn from a mess; models need a museum

Infants build robust vision and language from a firehose of blurry, first-person video and half-heard speech that rarely names what is on screen (Frank, 2023) (Long et al., 2025). VLMs do the opposite: they need web-scale, carefully curated image-caption pairs (Radford et al., 2021) (Alayrac et al., 2022), and still stumble on the egocentric streams that wearables and embodied agents produce. The reflex explanation is scale – children are just more sample-efficient. EgoBabyVLM (Lin et al., 2026) points at a different variable: the semantic alignment between what is seen and what is said. Control for alignment, and architecture and scale explain surprisingly little – while the data infants thrive on turns out to be barely more aligned than noise.

2. Making alignment measurable

A curated caption describes its image by construction. A toddler’s world does not: nearby speech – β€œcareful!”, β€œdo you want more?” – often refers to something off-screen, a moment ago, or nothing visible at all. To quantify this, embed matched (image, text) pairs with a pretrained CLIP encoder (Radford et al., 2021) (Bolya et al., 2025) and compare the distribution of cosine similarities against the same pairs with the captions shuffled. If the two distributions pull apart, the data is aligned; if they overlap, it is not. The size of that gap, in Jensen-Shannon divergence, is the alignment score (Hessel et al., 2021):

π’œοΈ€ = JSD ( 𝑃 βˆ₯ 𝑄 ) , 𝑃 ∼ { cos ( 𝑐 𝑖 , 𝑣 𝑖 ) } 𝑖 = 1 𝑁 , 𝑄 ∼ { cos ( 𝑐 πœ‹ ( 𝑖 ) , 𝑣 𝑖 ) } 𝑖 = 1 𝑁 ,

with πœ‹ a random permutation and π’œοΈ€βˆˆ[0,1]. Calibrating against deliberately-shuffled COCO gives a ruler from β€œaligned” down to β€œrandom,” and four corpora spread across it: curated captions land high, instructional video (Miech et al., 2019) in the middle, and naturalistic egocentric video – adult (Grauman et al., 2022) or infant (Long et al., 2025) – near the fully-shuffled floor (Tan et al., 2025). The corpus infants learn language from looks, to a pretrained encoder, close to noise.

3. The challenge, and a benchmark that grows with the model

EgoBabyVLM turns this into an experiment: train only on BabyView 2025.1 (about 863 hours of infant head-cam video), with no other data – not even to initialize an encoder (Long et al., 2025). Submissions are scored on cross-modal grounding, unimodal vision, and unimodal language, with the unimodal scores reported as a delta against a same-data baseline – so buying grounding by wrecking your own encoders is penalized.

Grounding is measured by Machine-DevBench, generated from the model’s own training vocabulary sampled across log-frequency bins (Algayres et al., 2025). This removes the vocabulary mismatch that cripples fixed benchmarks like DevBench (Tan et al., 2024) and BabyLM (Warstadt et al., 2023), where a baby-talk model gets quizzed on words it never heard. It yields Β 3,700 two-image, pick-the-match trials: two lexical (noun, adjective recognition) and eight grammatical (negation, word order, prepositions, counting, and so on).

4. Two findings

Two baselines bracket the design space: CLIP+ (contrastive, with interleaved InfoNCE, DINOv2 (Oquab et al., 2024), and masked-LM losses to resist forgetting) and LLaVA (generative captioning) (Liu et al., 2023).

1. Grounding tracks alignment, not architecture. Machine-DevBench accuracy rises monotonically with the training corpus’s alignment, for both backbones:

Training data Alignment Grounding (CLIP+)
BabyView (infant egocentric) weak β‰ˆ random 53.6
Ego4D (adult egocentric) weak β‰ˆ random 50.8
HowTo (instructional) moderate 56.1
COCO-MC (curated) high 66.5
off-the-shelf CLIP-L (reference) – 78.8

The egocentric corpora – the ones infants actually learn from – sit at chance (50%). The clincher is the shuffle series: hold model, size, and recipe fixed and only mismatch the captions, and grounding decays straight to chance.

% shuffled 0% 25% 50% 75% 100%
Grounding (CLIP+) 66.5 65.9 64.6 61.2 48.8

Nothing changed but alignment, so alignment is the causal factor – and BabyView’s failure is a property of the data as current objectives consume it, not a failure of the model.

2. Cross-modal training barely helps the unimodal encoders. The tide-lifts-all-boats hypothesis – seeing an object while hearing its name should sharpen both the picture and the word (Vong et al., 2024) – mostly fails here. Cross-modal finetuning never beats the same-data unimodal baseline on language: LLaVA loses roughly 10–14%, while CLIP+ stays within about 2% thanks to its interleaved unimodal losses. Even a probe built to detect visually-grounded semantics shows essentially no benefit.

5. Why it matters

Put together, a paradox: infants develop vision and language from BabyView-like input, yet models trained on it collapse to chance and gain nothing unimodally. The missing ingredients are likely inductive biases current VLMs lack – fine-grained patch-to-word attention, temporal integration, social cues like gaze and pointing, action grounding, and tolerance for weak alignment (Frank, 2023). By making alignment measurable and fixing the corpus behind an open leaderboard, EgoBabyVLM reframes the bottleneck: not a shortage of data or parameters (Villalobos et al., 2022), but a shortage of clean supervision – which infants plainly do not need (Sullivan et al., 2021) (Vong & Lake, 2025). The next generation of models will have to learn, as children do, from a world that rarely bothers to caption itself.

References

  1. Algayres, R., Saint-James, C.-Γ‰., Luthra, M., Shen, J., Benchekroun, Y., Lin, D., Moritz, R., Pino, J., and Dupoux, E. LongTail-Swap: benchmarking language models' abilities on rare words. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025.
  2. Long, B. L., Sparks, R. Z., Xiang, V., Stojanov, S., Yinzi, Keene, G., Tan, A. W. M., Feng, S. Y., Nag, A., Zhuang, C., Marchman, V. A., Yamins, D. L., and Frank, M. The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences. Proceedings of Cognitive Computational Neuroscience 2025, 2025.
  3. Sullivan, J., Mei, M., Perfors, A., Wojcik, E., and Frank, M. C. SAYCam: A Large, Longitudinal Audiovisual Dataset Recorded From the Infant's Perspective. Open Mind, 2021.
  4. Frank, M. C. Bridging the data gap between children and large language models. Trends in Cognitive Sciences, 2023.
  5. Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., and Choi, Y. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021.
  6. Tan, A. W. M., Yang, J., Sepuri, T., Aw, K. L., Sparks, R. Z., Yin, Z., Marchman, V. A., Frank, M. C., and Long, B. Assessing the alignment between infants' visual and linguistic experience using multimodal language models, 2025.
  7. Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual Instruction Tuning. Advances in Neural Information Processing Systems (Neurips), 2023.
  8. Vong, W. K., and Lake, B. M. On the robustness of modeling grounded word learning through a child's egocentric input, 2025.
  9. Warstadt, A., Mueller, A., Choshen, L., Wilcox, E., Zhuang, C., Ciro, J., Mosquera, R., Paranjabe, B., Williams, A., Linzen, T., and Cotterell, R. Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora. Proceedings of the Babylm Challenge at the 27th Conference on Computational Natural Language Learning, 2023.
  10. Lin, D., Rust, P., Corrales, A. V., Tan, A. W. M., Luthra, M., Saint-James, C.-Γ‰., Moritz, R., Krogh-Jespersen, S., Stark, V., Parimi, S., Shen, J., Benchekroun, Y., Higuchi, Y., Gleize, M., Fizycki, T., Hamilakis, N., Khentout, M., Tsuji, S., KΓ©gl, B., … Dupoux, E. EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data, 2026.
  11. Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., Martin, M., Nagarajan, T., Radosavovic, I., Ramakrishnan, S. K., Ryan, F., Sharma, J., Wray, M., Xu, M., Xu, E. Z., … Malik, J. Ego4D: Around the World in 3,000 Hours of Egocentric Video. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  12. Oquab, M., Darcet, T., Moutakanni, T., Vo, H. V., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., … Bojanowski, P. DINOv2: Learning Robust Visual Features without Supervision. Transactions on Machine Learning Research (TMLR), 2024.
  13. Vong, W. K., Wang, W., Orhan, A. E., and Lake, B. M. Grounded language acquisition through the eyes and ears of a single child. Science, 2024.
  14. Villalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L., and Hobbhahn, M. Will we run out of data? Limits of LLM scaling based on human-generated data, 2022.
  15. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning Transferable Visual Models From Natural Language Supervision. International Conference on Machine Learning (ICML), 2021.
  16. Bolya, D., Huang, P.-Y., Sun, P., Cho, J. H., Madotto, A., Wei, C., Ma, T., Zhi, J., Rajasegaran, J., Rasheed, H., Wang, J., Monteiro, M., Xu, H., Dong, S., Ravi, N., Li, D., DollΓ‘r, P., and Feichtenhofer, C. Perception Encoder: The best visual embeddings are not at the output of the network, 2025.
  17. Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J. L., Borgeaud, S., … Simonyan, K. Flamingo: A Visual Language Model for Few-Shot Learning. Advances in Neural Information Processing Systems (Neurips), 2022.
  18. Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., and Sivic, J. HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. IEEE/CVF International Conference on Computer Vision (ICCV), 2019.
  19. Tan, A. W., Yu, S., Long, B., Ma, W. A., Murray, T., Silverman, R. D., Yeatman, J. D., and Frank, M. C. DevBench: A multimodal developmental benchmark for language learning. Advances in Neural Information Processing Systems (Neurips), 2024.