From Pixels to Semantics: Where Streaming Trends Meet the Container

From Pixels to Semantics: Where Streaming Trends Meet the Container

ITU-T SG21 and ISO/IEC JTC 1 SC 29 Joint Workshop on “Media Streaming Service – What’s next”

Geneva, July 14, 2026

[Workshop][Slides][PDF]

Christian Timmerer (AAU/Bitmovin, Austria)

Abstract: The next inflection in streaming isn’t a codec — it’s three pressures (efficiency, low-latency live, and AI-native media) converging on the systems layer, with concrete implications for the MPEG Systems and ITU-T SG21 roadmap drawn from Bitmovin and the ATHENA lab.

Posted in ATHENA | Comments Off on From Pixels to Semantics: Where Streaming Trends Meet the Container

LumaID: Harnessing Illumination-Awareness for High-Fidelity Video Head Identity Editing

LumaID: Harnessing Illumination-Awareness for High-Fidelity Video Head Identity Editing

34th ACM International Conference on Multimedia 2026 (ACM MM 2026)

10–14 November 2026

Rio de Janeiro, Brazil

[PDF]

Yiying Wei (AAU, Austria), Xuanhong Chen (Shanghai Jiao Tong University, China), Hadi Amirpour (AAU, Austria) and Christian Timmerer (AAU, Austria)

Abstract: Despite yielding higher visual quality than image-to-image approaches, masked generation paradigms for video face editing fundamentally lacks attribute consistency (e.g., illumination, background). We introduce LumaID, a novel framework that explicitly disentangles identity and expression representations from environmental contexts, enabling high-fidelity, fine-grained video head editing while strictly preserving these crucial attributes. At its core, LumaID employs an Omni-Disentangled Diffusion Transformer (OD-DiT) that leverages 3D proxy representations to thoroughly isolate the source and target facial features, fundamentally preventing identity leakage and illumination degradation. To further overcome the distributional drift caused by proxy estimation noise and the lack of explicit consistency supervision, we propose Consist-GRPO. This post-training reinforcement learning mechanism formulates multi-dimensional reward signals (spanning identity, expression, pose, and lighting) to continuously steer the generative process toward strict spatiotemporal alignment. Extensive evaluations demonstrate that LumaID serves as a highly competitive baseline, exhibiting strong performance over prior approaches in both attributes consistency and overall visual quality.

Posted in ATHENA | Comments Off on LumaID: Harnessing Illumination-Awareness for High-Fidelity Video Head Identity Editing

Thin-Client Interactive Gaussian Adaptive Streaming over HTTP/3 (accepted in ACM MM’26)

Thin-Client Interactive Gaussian Adaptive Streaming over HTTP/3

ACM Multimedia 2026

November 10 – November 14, 2026

Rio de Janeiro, Brazil

[PDF]

Emanuele Artioli (AAU, Austria), Philipp Fößl (AAU, Austria), Shao-Yang Hung (National Tsinghua University, Taiwan), Philipp Fößl (AAU, Austria), Daniele Lorenzi (Bitmovin, Austria), Farzad Tashtarian (AAU, Austria),  Mahdi Dolati (Sharif University of Technology, Iran), Cheng-Hsin Hsu (National Tsinghua University, Taiwan), Christian Timmerer (AAU, Austria)

Abstract: Recent advancements in 3D Gaussian Splatting (3DGS) have enabled photorealistic rendering of complex scenes, yet widespread adoption on mobile and Extended Reality (XR) devices is hindered by substantial computational and bandwidth requirements. While existing solutions often focus on model compression for client-side rendering, they still demand significant GPU power, limiting applicability on resource-constrained hardware. We propose TIGAS (Thin-client Interactive Gaussian Adaptive Streaming), a remote rendering framework offloading rasterization to a backend. To bypass the prohibitive latencies connected to fluctuating network conditions, TIGAS streams view-dependent 2D projections to a lightweight web client over QUIC, minimizing head-of-line (HoL) blocking. A dedicated ABR algorithm adapts rendering quality to fluctuating network conditions, maintaining motion-to-photon latency within strict 6DoF interactive constraints. Furthermore, we discuss the integration of an experimental WebGPU super-resolution pipeline to analyze the trade-offs between perceptual quality enhancements and thin-client processing bottlenecks. We extensively evaluate TIGAS across multi-continental environments using 14 3DGS models and real 6DoF EyeNavGS movement traces. Powered by a backend rendering frames in under 10 milliseconds, TIGAS maintains latency within interactive thresholds while achieving an average SSIM of 0.88, serving both as a robust testbed for 3DGS streaming research and a capable delivery system.

Posted in ATHENA | Comments Off on Thin-Client Interactive Gaussian Adaptive Streaming over HTTP/3 (accepted in ACM MM’26)

LMM-10K: Large-Scale 4K Multimodal Dataset for Perceptual, Semantic, and Content-Aware Video Processing

LMM-10K: Large-Scale 4K Multimodal Dataset for Perceptual, Semantic, and Content-Aware Video Processing

ACM Multimedia 2026

November 10 – November 14, 2026

Rio de Janeiro, Brazil

[PDF]

Mohammad Ghasempour (AAU, Austria), Yiying Wei (AAU, Austria), Hadi Amirpour (AAU, Austria), Christian Timmerer (AAU, Austria)

Abstract: The growing integration of vision and language models is driving a fundamental shift in video understanding and processing. This evolution calls for datasets that jointly capture visual content and its semantic representations at scale. To address this need, we introduce LMM-10K, a large-scale, curated multimodal dataset comprising 10,000 high-fidelity 4K video sequences at 60 fps with rich semantic and perceptual annotations. We developed an automated acquisition pipeline to curate videos from the Pexels repository, using targeted search queries and strict filtering criteria to capture a wide range of real-world scenes. Beyond the video sequences, LMM-10K is enriched with comprehensive multimodal annotations that integrate low-level visual features with high-level semantic information. These include LLM-generated semantic descriptors, no-reference quality metrics, spatial-temporal complexity metrics, and visual diversity attributes. By combining structured annotations with high-quality video data, LMM-10K provides a versatile resource for a wide range of applications, including video enhancement, content-aware compression and streaming, neural video coding, multimodal learning, generative video modeling, and perceptual quality modeling. Dataset URL: Link

Posted in ATHENA | Comments Off on LMM-10K: Large-Scale 4K Multimodal Dataset for Perceptual, Semantic, and Content-Aware Video Processing

Selective Multi-Pass Encoding for Cost-Efficient Video Streaming

Selective Multi-Pass Encoding for Cost-Efficient Video Streaming

International Broadcasting Convention (IBC)

[PDF]

Mohammad Ghasempour (AAU, Austria), Hadi Amirpour (AAU, Austria), and Christian Timmerer (AAU, Austria)

Abstract: As video streaming scales across platforms, resolutions, and devices, encoding efficiency has become critical to maintaining quality while controlling computational cost and energy consumption. Multi-pass encoding is widely used in streaming workflows to improve compression efficiency, rate-control accuracy, and quality consistency. However, its computational overhead is applied uniformly across all content, even when additional passes deliver minimal benefit. At scale, this results in unnecessary processing, higher computational cost, and increased energy consumption. This paper presents a content-adaptive strategy, CASE, that predicts whether additional encoding passes would provide meaningful gains using a lightweight mechanism that derives spatial and temporal features from each video segment. We evaluated the approach in two production-oriented scenarios using local and cloud-based video encoders. Results show that the method reduces computational time and encoding cost, with minimal impact on compression efficiency and visual quality. Experimental results show that CASE reduces encoding time by 25.3% on average with only a 2.23% bitrate increase, while its preprocessing and decision overhead is about 1155 times lower than multi-pass encoding time.

Posted in ATHENA | Comments Off on Selective Multi-Pass Encoding for Cost-Efficient Video Streaming

The 4th Workshop on Emerging Multimedia Systems (EMS) 2026

The 4th Workshop on Emerging Multimedia Systems (EMS) 2026

 EMS 2023  |  EMS 2024  | EMS 2025

Multimedia has played a significant role in driving Internet usage and has led to a range of technological advancements, such as content delivery networks, compression algorithms, and streaming protocols. With emerging applications, including (but not limited to) augmented, virtual, and extended reality (XR), real-time telepresence, AI-generated content, video analytics, and the usage of AI in multimedia systems in general, multimedia is undergoing a fundamental shift in sharing experiences online and continues to drive the future of the Internet. As these next-generation ultra-low-latency, interactive, and immersive technologies evolve, it is crucial to revisit developed techniques for new formats and representations, not only to enhance performance and interactivity but also to improve energy efficiency and maintain high Quality of Experience (QoE). This workshop will bring together experts from diverse fields, including video streaming research, source video coding, analytics, rate adaptation algorithms, networked systems, immersive media such as 3D and volumetric video streaming, AR/VR applications, as well as energy-efficient systems and QoE optimization, to exchange ideas on identifying challenges and opportunities in designing advanced networked systems for these emerging multimedia technologies (more details)

Posted in ATHENA | Comments Off on The 4th Workshop on Emerging Multimedia Systems (EMS) 2026

Cross-Layer Dynamics in Live Low-Latency: A Dataset of ABR, CC, and AQM Interactions

Cross-Layer Dynamics in Live Low-Latency: A Dataset of ABR, CC, and AQM Interactions

18th International Conference on Quality of Multimedia Experience

Cardiff, UK, June 29th – July 3rd, 2026

[PDF]

Md Tariqul Islam (UNICAMP, Brazil),  Farzad Tashtarian (AAU, Austria),  Christian Esteve Rothenberg (UNICAMP, Brazil), Christian Timmerer (AAU, Austria)

Low-latency video streaming, such as Low-Latency DASH (LL-DASH), requires maintaining high Quality of Experience (QoE) under varying network conditions. In LL-DASH, QoE is jointly influenced not only by Adaptive Bitrate (ABR) decisions, but also by transport-layer Congestion Control (CC) and network-layer Active Queue Management (AQM), whose interactions remain insufficiently characterized due to limited cross-layer experimentation. Therefore, we present a large-scale LL-DASH dataset comprising approximately 2,000 controlled sessions across three dash.js ABR algorithms (L2A, Dynamic, LoLP), three CC schemes (CUBIC, BBRv1, Prague) across both TCP and QUIC transport protocols, four AQM configurations (FIFO, FQ-CoDel, CAKE, DualPI2), and multiple congestion scenarios. The dataset supports QoE-aware cross-layer analysis and ABR benchmarking under diverse network configurations and is available at: https://github.com/cd-athena/ ll-dash-crosslayer-dataset

Posted in ATHENA | Comments Off on Cross-Layer Dynamics in Live Low-Latency: A Dataset of ABR, CC, and AQM Interactions