Depth-Aware Stereo Bitrate Ladder Optimization with Decoding-Time Constraints for HTTP Adaptive Streaming

Depth-Aware Stereo Bitrate Ladder Optimization with Decoding-Time Constraints for HTTP Adaptive Streaming

IEEE International Workshop on Multimedia Signal Processing (MMSP)

September 22 – September 24, 2026

Istanbul, Turkey

[PDF]

Kamran Qureshi (AAU, Austria), Hadi Amirpour (AAU, Austria), Farzad Tashtarian (AAU, Austria), Christian Timmerer (AAU, Austria)

Abstract: Per-title bitrate ladder construction selects bitrate–resolution pairs based on content-specific characteristics, enabling improved compression efficiency compared to static bitrate ladders. Extending content-adaptive bitrate ladder construction to stereoscopic video, we propose a content- and depth-aware stereoscopic bitrate ladder that jointly optimizes three dimensions: (i) perceptual video quality, (ii) depth fidelity, and (iii) decoding efficiency within a unified optimization framework using objective quality metrics Advanced Video Quality Tool (AVQT), Just Noticeable Difference in Depth (JNDD)-filtered depth violations, and decoding time measurements. Bitrate ladder construction is formulated as a binary linear programming optimization problem that selects one representation at each bitrate subject to constraints, with tunable weighting to balance the three objectives. Experimental results demonstrate that the proposed approach achieves a balanced trade-off across perceptual quality, depth fidelity, and decoding efficiency, yielding, on average, a 4.61% BD-rate reduction in perceptual quality, a 2.07% reduction in depth violations, and a 10.49% reduction in decoding time relative to a baseline comprising a fixed bitrate ladder. The resulting bitrate ladders consistently outperform both baselines that optimize individual objectives in isolation and the fixed bitrate ladder, highlighting the benefits of jointly optimizing quality, depth, and decoding efficiency for stereoscopic video streaming.

Posted in ATHENA | Comments Off on Depth-Aware Stereo Bitrate Ladder Optimization with Decoding-Time Constraints for HTTP Adaptive Streaming

Viewport-Aware Adaptive Encoding for Enhanced 360-degree Video Streaming

Viewport-Aware Adaptive Encoding for Enhanced 360-degree Video Streaming

International Workshop on Multimedia Signal Processing (MMSP)

September 22 – September 24, 2026

Istanbul, Türkiye

[PDF]

Mahmoud Z. A. Wahba (University of Padova), Mohammad Ghasempour (AAU, Austria), Sara Baldoni (University of Padova), Christian Timmerer (AAU, Austria), Federica Battisti (University of Padova), Hadi Amirpour (AAU, Austria)

Abstract: Streaming 360-degree video content in Virtual Reality (VR) poses significant challenges, particularly in balancing perceived quality and available bandwidth. Tile-based streaming guided by viewport prediction reduces bandwidth usage by allocating a higher bitrate to tiles within the predicted viewport and a lower bitrate to non-viewport tiles. However, viewport prediction is not always precise, and failure cases can significantly degrade the user’s Quality of Experience. In this paper, we address this challenge by selecting the optimal encoding configuration, i.e., the video encoding resolution, for each tile to improve the quality of tiles outside the predicted viewport without increasing the video target bitrate. We evaluate our method on a large-scale 360-degree video dataset, which demonstrates our model effectiveness in improving user-perceived quality, particularly when viewport prediction fails. Our experimental results show a 37.41% improvement in BD-Rate and a 7.5dB gain in BD-PSNR for the viewport-mispredicted tiles.

Posted in ATHENA | Comments Off on Viewport-Aware Adaptive Encoding for Enhanced 360-degree Video Streaming

Content-Adaptive Encoding Pass Selection for Efficient Video Streaming

Content-Adaptive Encoding Pass Selection for Efficient Video Streaming

International Workshop on Multimedia Signal Processing (MMSP)

September 22 – September 24, 2026

Istanbul, Türkiye

[PDF]

Mohammad Ghasempour (AAU, Austria), Hadi Amirpour (AAU, Austria), Christian Timmerer (AAU, Austria)

Abstract: As video streaming continues to grow in scale, improving the efficiency of video encoding has become increasingly important to balance visual quality and computational cost. Multi-pass encoding is widely adopted to enhance compression efficiency, improve rate control, and achieve more stable quality by leveraging additional analysis of video content prior to encoding. These benefits come with the cost of increased computational complexity. In this paper, we show that the benefits of multi-pass encoding vary substantially across video content and encoding configurations in adaptive video streaming. Motivated by this observation, we propose the Adaptive Encoding Pass Selection (AEPS), a lightweight content-adaptive framework that estimates the benefits of multi-pass prior to encoding and enables selective use of single-pass encoding to reduce encoding time. Experimental results demonstrate that the AEPS framework substantially reduces encoding time while maintaining compression performance and stability, achieving an average 25.3% reduction in encoding time with only a 2.23% increase in bitrate. We show that the preprocessing and decision-making overhead of AEPS is approximately 1155 times lower than the time required for multi-pass encoding in adaptive streaming.

Posted in ATHENA | Comments Off on Content-Adaptive Encoding Pass Selection for Efficient Video Streaming

From Pixels to Semantics: Where Streaming Trends Meet the Container

From Pixels to Semantics: Where Streaming Trends Meet the Container

ITU-T SG21 and ISO/IEC JTC 1 SC 29 Joint Workshop on “Media Streaming Service – What’s next”

Geneva, July 14, 2026

[Workshop][Slides][PDF]

Christian Timmerer (AAU/Bitmovin, Austria)

Abstract: The next inflection in streaming isn’t a codec — it’s three pressures (efficiency, low-latency live, and AI-native media) converging on the systems layer, with concrete implications for the MPEG Systems and ITU-T SG21 roadmap drawn from Bitmovin and the ATHENA lab.

Posted in ATHENA | Comments Off on From Pixels to Semantics: Where Streaming Trends Meet the Container

LumaID: Harnessing Illumination-Awareness for High-Fidelity Video Head Identity Editing

LumaID: Harnessing Illumination-Awareness for High-Fidelity Video Head Identity Editing

34th ACM International Conference on Multimedia 2026 (ACM MM 2026)

10–14 November 2026

Rio de Janeiro, Brazil

[PDF]

Yiying Wei (AAU, Austria), Xuanhong Chen (Shanghai Jiao Tong University, China), Hadi Amirpour (AAU, Austria) and Christian Timmerer (AAU, Austria)

Abstract: Despite yielding higher visual quality than image-to-image approaches, masked generation paradigms for video face editing fundamentally lacks attribute consistency (e.g., illumination, background). We introduce LumaID, a novel framework that explicitly disentangles identity and expression representations from environmental contexts, enabling high-fidelity, fine-grained video head editing while strictly preserving these crucial attributes. At its core, LumaID employs an Omni-Disentangled Diffusion Transformer (OD-DiT) that leverages 3D proxy representations to thoroughly isolate the source and target facial features, fundamentally preventing identity leakage and illumination degradation. To further overcome the distributional drift caused by proxy estimation noise and the lack of explicit consistency supervision, we propose Consist-GRPO. This post-training reinforcement learning mechanism formulates multi-dimensional reward signals (spanning identity, expression, pose, and lighting) to continuously steer the generative process toward strict spatiotemporal alignment. Extensive evaluations demonstrate that LumaID serves as a highly competitive baseline, exhibiting strong performance over prior approaches in both attributes consistency and overall visual quality.

Posted in ATHENA | Comments Off on LumaID: Harnessing Illumination-Awareness for High-Fidelity Video Head Identity Editing

Thin-Client Interactive Gaussian Adaptive Streaming over HTTP/3 (accepted in ACM MM’26)

Thin-Client Interactive Gaussian Adaptive Streaming over HTTP/3

ACM Multimedia 2026

November 10 – November 14, 2026

Rio de Janeiro, Brazil

[PDF]

Emanuele Artioli (AAU, Austria), Philipp Fößl (AAU, Austria), Shao-Yang Hung (National Tsinghua University, Taiwan), Philipp Fößl (AAU, Austria), Daniele Lorenzi (Bitmovin, Austria), Farzad Tashtarian (AAU, Austria),  Mahdi Dolati (Sharif University of Technology, Iran), Cheng-Hsin Hsu (National Tsinghua University, Taiwan), Christian Timmerer (AAU, Austria)

Abstract: Recent advancements in 3D Gaussian Splatting (3DGS) have enabled photorealistic rendering of complex scenes, yet widespread adoption on mobile and Extended Reality (XR) devices is hindered by substantial computational and bandwidth requirements. While existing solutions often focus on model compression for client-side rendering, they still demand significant GPU power, limiting applicability on resource-constrained hardware. We propose TIGAS (Thin-client Interactive Gaussian Adaptive Streaming), a remote rendering framework offloading rasterization to a backend. To bypass the prohibitive latencies connected to fluctuating network conditions, TIGAS streams view-dependent 2D projections to a lightweight web client over QUIC, minimizing head-of-line (HoL) blocking. A dedicated ABR algorithm adapts rendering quality to fluctuating network conditions, maintaining motion-to-photon latency within strict 6DoF interactive constraints. Furthermore, we discuss the integration of an experimental WebGPU super-resolution pipeline to analyze the trade-offs between perceptual quality enhancements and thin-client processing bottlenecks. We extensively evaluate TIGAS across multi-continental environments using 14 3DGS models and real 6DoF EyeNavGS movement traces. Powered by a backend rendering frames in under 10 milliseconds, TIGAS maintains latency within interactive thresholds while achieving an average SSIM of 0.88, serving both as a robust testbed for 3DGS streaming research and a capable delivery system.

Posted in ATHENA | Comments Off on Thin-Client Interactive Gaussian Adaptive Streaming over HTTP/3 (accepted in ACM MM’26)

LMM-10K: Large-Scale 4K Multimodal Dataset for Perceptual, Semantic, and Content-Aware Video Processing

LMM-10K: Large-Scale 4K Multimodal Dataset for Perceptual, Semantic, and Content-Aware Video Processing

ACM Multimedia 2026

November 10 – November 14, 2026

Rio de Janeiro, Brazil

[PDF]

Mohammad Ghasempour (AAU, Austria), Yiying Wei (AAU, Austria), Hadi Amirpour (AAU, Austria), Christian Timmerer (AAU, Austria)

Abstract: The growing integration of vision and language models is driving a fundamental shift in video understanding and processing. This evolution calls for datasets that jointly capture visual content and its semantic representations at scale. To address this need, we introduce LMM-10K, a large-scale, curated multimodal dataset comprising 10,000 high-fidelity 4K video sequences at 60 fps with rich semantic and perceptual annotations. We developed an automated acquisition pipeline to curate videos from the Pexels repository, using targeted search queries and strict filtering criteria to capture a wide range of real-world scenes. Beyond the video sequences, LMM-10K is enriched with comprehensive multimodal annotations that integrate low-level visual features with high-level semantic information. These include LLM-generated semantic descriptors, no-reference quality metrics, spatial-temporal complexity metrics, and visual diversity attributes. By combining structured annotations with high-quality video data, LMM-10K provides a versatile resource for a wide range of applications, including video enhancement, content-aware compression and streaming, neural video coding, multimodal learning, generative video modeling, and perceptual quality modeling. Dataset URL: Link

Posted in ATHENA | Comments Off on LMM-10K: Large-Scale 4K Multimodal Dataset for Perceptual, Semantic, and Content-Aware Video Processing