Efficient Multimodal Question Answering (EMM-QA)

π List of Accepted Papers and Shared Challenge Papers
EMM-QA is an ICML 2026 workshop focused on question answering systems that must balance accuracy, efficiency, and adaptability across multiple input modalities. The workshop brings together researchers from academia and industry working on knowledge-intensive multimodal systems that operate under practical resource constraints.
Rather than focusing only on larger models, the workshop emphasizes methods that make multimodal question answering usable in real settings, including retrieval-augmented systems, compact models, efficient inference, and human-in-the-loop evaluation.
This yearβs shared challenge explored that idea in a game setting: players and systems answered pyramid-style questions that mixed text with images, then adapted through tossups, bonuses, confidence, and human-computer teaming. We also used adversarial question writing to probe where multimodal models are still brittle, especially on surprising clues that are easy for humans but hard for systems.
For the human-computer matches, the AI teammates shown in the presentation as a grid of model logos were drafted into teams in ranking order: weaker human teams chose first, then the draft continued upward through the standings and reversed if AI teammates were still available. That pairing of roster and draft order is what makes the teammate selection visually meaningful.
QANTA 2026 challenge CSVs are available for Google Sheets: tossup question difficulty and bonus question difficulty.
Scope
The workshop is centered on efficient multimodal question answering. It also welcomes closely related work on multimodal retrieval, reasoning, evaluation, benchmarking, and efficient inference when those contributions are clearly connected to question answering or other knowledge-intensive multimodal tasks.
Like the previous iteration of EfficientQA, which focused on text-only question answering, we will also host a human-computer question answering competition. If youβd like to take part in that part of the competition (it should be fun!), you can either play as a team or write questions.
Workshop Format
The workshop is planned as a one-day event combining:
- Contributed papers
- Poster presentations
- Invited keynotes
- Shared-task highlights
- A live human-computer question answering event
The workshop will also serve as the venue where we announce the winning systems from the QANTA 2026 computer competition.
The presentation also closes with a few reflections: multimodal systems remain strong overall, but they still struggle with surprising visual clues, explanation quality, calibration, and adversarial generation. That makes the competition useful not just as a leaderboard, but as a way to study the jagged frontier between human and machine performance.
Schedule
- Workshop takes place on July 11th. Both poster sessions
- All talk sessions (invited talks, spotlights, challenge talks, awards, etc.) will take place in the ASEM Ballroom 201 at COEX.
- All poster sessions will take place separately in Hall A, poster boards 1612β1617, 1700β1711, outside the workshop room area at COEX. You can review the Hall A plan here.
| Time | Activity | Duration |
|---|---|---|
| 08:00β08:10 | Welcome & Workshop Overview | 10 min |
| 08:10β08:50 | π¦ Naman Goyal & Jenny Ni: Multimodal Robustness Under Distribution Shift | 40 min |
| 08:50β09:00 | Q&A | 10 min |
| 09:00β09:15 | β Coffee Break | 15 min |
| 09:15β09:55 | π¦ Sewon Min: PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation | 40 min |
| 09:55β10:05 | Q&A | 10 min |
| 10:05β10:50 | π¨ Contributed Paper Spotlights I | 45 min |
| 10:05β10:15 | π¨ Mirage Probes: How Vision Models Fake Visual Understanding β Daniel Ben-Levi et al. | 10 min |
| 10:16β10:26 | π¨ VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following β Hyesoo Hong et al. | 10 min |
| 10:27β10:37 | π¨ DistortBench: Benchmarking Vision Language Models on Image Distortion Identification β Divyanshu Goyal et al. | 10 min |
| 10:38β10:48 | π¨ CCDiff: Inverse Canonical Correlation Analysis for Discovering Visual Differences in Natural Language (Virtual) β Neelesh Bisht et al. | 10 min |
| 10:50β11:50 | π§ Workshop Posters I | 60 min |
| 11:50β12:50 | Lunch | 60 min |
| 12:50β13:20 | π€ Live AI QA Competition | 30 min |
| 13:20β14:00 | π¦ Mrinmaya Sachan: Behavioral Shortcuts and Cross-Modal Alignment Failures of Multimodal Large Language Models | 40 min |
| 14:00β14:10 | Q&A | 10 min |
| 14:10β14:50 | π¦ Robin Jia: The Golden Age of Adversarial Evaluation | 40 min |
| 14:50β15:00 | Q&A | 10 min |
| 15:00β15:15 | β Coffee Break | 15 min |
| 15:15β15:35 | Shared Challenge Introduction & Results Overview | 20 min |
| 15:35β15:55 | π¨ Contributed Paper Spotlights II | 20 min |
| 15:35β15:45 | π¨ Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment β Harikrishnan Puthan Madathil et al. | 10 min |
| 15:45β15:55 | π¨ FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models (Virtual) β Youngsun Lim et al. | 10 min |
| 15:55β16:05 | π Challenge Awards | 10 min |
| 16:05β16:10 | Closing Remarks | 5 min |
| 16:10β17:00 | π§ Workshop Posters II + Shared Challenge Posters | 50 min |
Legend
- π¦ Invited Talks
- π¨ Contributed Paper Spotlights / Best Challenge Team Talk
- π§ Poster Sessions
Sponsors/Acknowledgements
- This workshop is partially supported by Horizon EU programme through project ELOQUENCE, grant no. 101135916.