Concept detail
multimodal conditioning
0 tagged videos · 1 connection videos · degree 0
Publication timeline
- 2026-081 video
Source videos
Observed neighbors
language-as-lossy-compression
text is a bottleneck for sensory fidelity, so reference images/video/audio must condition the model directly to bridge user intent and output.
1 source videos
SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMindRelated tagged insights
No core insights are tagged with this concept.