Two papers have been accepted to EMNLP 2026 (Main) 
This year, EMNLP received an unprecedented 17,669 submissions. The Program Committee accepted 2,719 papers to the Main Conference and 2,533 papers to Findings, resulting in a Main Conference acceptance rate of 15.4%.
Title: IGG: A Benchmark for Interactive GUI Grounding under Visibility Constraints
Authors: KyeongSeon Kim* (KAIST), Jiyeon Son* (KAIST), Tae-Hyun Oh (KAIST)
GUI grounding benchmarks usually ask whether a VLM agent can localize a target from a static screenshot where the target is already visible. Real interfaces often violate this assumption: targets may be off-screen, hidden behind UI state, occluded, ambiguous until hover, or activated only after a delay. We introduce Interactive GUI Grounding (IGG), a benchmark for grounding under limited observability. In IGG, agents must first expose a hidden or non-actionable target, then click it. IGG defines a minimal action space and seven visibility-constraint sub-types spanning viewport constraint, multi-state constraint, and temporal and manipulation-based settings. Across 16 controllable mock GUI environments and 601 tasks, humans achieve 94.6% grounding accuracy, while strong open-source VLM agents achieve only around 30%, revealing a large gap in visibility recovery rather than ordinary target localization.
Title: Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
Authors: Oh Hyun-Bin (POSTECH), Kazuki Shimada (Sony), Yuhta Takida (Sony), Kim Sung-Bin (POSTECH), Toshimitsu Uesaka (Sony), Takashi Shibuya (Sony), Kyeongyoon Lee (SKKU), Tae-Hyun Oh (KAIST)†, Yuki Mitsufuji (Sony)†
(†Co-corresponding authors)
Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track source directions over time but offer limited semantic coverage for language reasoning. To address this gap, we introduce ST-AudioQA, a spatio-temporal audio QA dataset and benchmark built from first-order ambisonic (FOA) renderings of static and moving sound sources. Each scene provides source identity, activity, direction, distance, and motion metadata, enabling dense trajectory supervision and questions about what is sounding, where it is, how it moves, and how sources relate. We further propose ST-Audio Encoder, a time-resolved FOA audio encoder that learns event semantics together with source trajectories, and ST-AudioLM, which connects the audio tokens from the encoder to an LLM for spatio-temporal audio QA. Experiments show that this representation improves the semantic-localization tradeoff and yields stronger reasoning performance than static spatial and localization-oriented baselines.




