Efficient Quality Controller for Video Encoding

Efficient Quality Controller for Video Encoding

IEEE Visual Communications and Image Processing Conference (VCIP 2026)

December 13–16, 2026

Singapore

[PDF]

Yiying Wei (AAU, Austria), Hadi Amirpour (AAU, Austria),  and Christian Timmerer (AAU, Austria)

Abstract: Traditional video streaming relies on Adaptive Bitrate (ABR) algorithms that encode videos at fixed bitrate-resolution pairs. As a result, a rate controller is essential to ensure that each encoded representation meets its target bitrate. However, perceptually-aware bitrate ladder construction methods aim to encode videos at a fixed visual quality instead of a fixed bitrate, to avoid under- or over-allocating bits for complex and simple content. In this paper, we propose an efficient quality controller that predicts the Quantization Parameter (QP) required to achieve a target VMAF score for each video segment. The framework supports both CPU-only operation for low-complexity environments and GPU-accelerated inference for improved prediction accuracy. By leveraging content features and target quality levels, our model estimates appropriate QP values without requiring pre-encoding or tight integration with the encoder. For target VMAF scores of 94, 88, and 82, the CPU-only model achieves mean absolute errors (MAEs) of 1.05, 1.24, and 1.34, respectively, comparable to the state-of-the-art errors of 1.14, 1.27, and 1.31, while requiring only a fraction of the computational cost. The GPU-based model further reduces the MAEs to 0.50, 0.49, and 0.47, less than half of the state-of-the-art errors.

Posted in ATHENA | Comments Off on Efficient Quality Controller for Video Encoding

Cloud, Edge, or Split? Profiling Onboard and Split Vision-Language Model Deployment for Drone AI

Cloud, Edge, or Split? Profiling Onboard and Split Vision-Language Model Deployment for Drone AI

Zoha Azimi, Reza Farahani, Schahram Dustdar, Christian Timmerer

The 4th International Symposium on Edge Intelligence, Trustworthy and Decentralized Artificial Intelligence (iEDGE 2026)

October 27-30, 2026 – Paris, France

Vision-Language Models (VLMs) enable edge devices like unmanned aerial vehicles (UAVs) to interpret visual observations and reason about complex environments using natural-language instructions. However, their practical deployment remains challenging as onboard inference is constrained by limited computational, memory, and energy resources, whereas cloud-based inference introduces communication latency, bandwidth overhead, and dependence on network connectivity. To address these limitations, split computing offers a promising alternative by partitioning VLM inference between the resource-constrained UAVs and more capable remote servers. However, the performance trade-offs among fully onboard, cloud-based, and split-computing architectures for lightweight VLMs have not yet been systematically profiled. This paper benchmarks these three deployment paradigms using SmolVLM-256M as a representative lightweight VLM. We quantify their inference latency, computational resource utilization, communication overhead, and energy consumption across varying image resolutions and network conditions. Our results show that no deployment strategy is universally optimal; instead, the preferred strategy depends on the interaction between network conditions and input image resolution.

Posted in ATHENA | Comments Off on Cloud, Edge, or Split? Profiling Onboard and Split Vision-Language Model Deployment for Drone AI

ATHENA Closing Symposium

Seven Years of Adaptive Streaming Research and What Comes Next

Join us as we celebrate the conclusion of the Christian Doppler Laboratory ATHENA [PDF)].

After seven years of research into adaptive video streaming and emerging networked multimedia services, we will look back at ATHENA’s achievements and explore the technologies, collaborations, and research questions shaping the future of video streaming.

Wednesday, 7 October 2026, 14:00–17:00
Alpen-Adria-Universität Klagenfurt, Stiftungssaal · Room O.0.0.1

Programme

14:00–14:05 · Welcome
14:05–14:35 · KeynoteTime to MOQ On: Leaving Legacy Latency Behind? (Ali C. Begen, Özyeğin University)
14:35–14:45 · Seven Years of ATHENA: Achievements and Impact (Christian Timmerer)
14:45–15:00 · Coffee Break
15:00–15:15 · Network-Assisted Adaptive Streaming: Toward Optimal QoE through System Collaboration (Farzad Tashtarian)
15:15–15:30 · Beyond One-Size-Fits-All: Adaptive and Computationally Efficient Video Streaming (Hadi Amirpour)
15:30–16:00 · Industry Panel: The Future of Video Streaming Chaired by Reinhard Grandl, Chief Product Officer at Bitmovin with Bitmovers and invited guests
From 16:00 · Drinks, Bites and Networking

We look forward to celebrating seven years of ATHENA with colleagues, collaborators, and friends — and to continuing the conversation about what comes next.

Free admission · Registration required via itec-sek@itec.aau.at

Posted in ATHENA | Comments Off on ATHENA Closing Symposium

Interns at ATHENA (Summer 2026)

Interns at ATHENA (Summer 2026)

In July 2026, the ATHENA Christian Doppler Laboratory hosted four interns working on the following topics:

  • Leon Kordasch – Holography
  • Daniel Glantschnig – Automated Wind Turbine Damage Detection
  • Gabriel Puri – Adaptive Streaming for Immersive Media

At the end of their internships, the interns presented their projects and findings and received official university certificates in recognition of their work. The experience proved valuable for both the interns and the ATHENA research team alike. Through personalized mentorship, hands-on training, and continuous support, the interns were able to develop strong practical skills while gaining a deeper understanding of research methodologies and technologies in the video streaming domain. We warmly thank the interns for their enthusiasm, dedication, and thoughtful feedback, which made a meaningful contribution to the ongoing work of the ATHENA lab.

Leon Kordasch: “During my internship, I worked on digital holography. I explored state-of-the-art solutions, analyzed their performance and gained lots of theoretical knowledge and technical experience. While challenging, the internship was very rewarding. My supervisor, Ayman Alkhateeb, provided guidance where needed, and collaborating with a diverse, international team made the experience both enriching and enjoyable.”

Daniel Glantschnig: “My time as an intern was both interesting and rewarding. I had the opportunity to train object detection models using datasets with and without synthetic data, then compare the results to explore how synthetic data influenced model performance. Working on this project helped me gain a much better understanding of dataset preparation, model training, evaluation, and the impact that different types of data can have on object detection systems. One of the highlights of the internship was the welcoming and friendly team, which made the experience even more enjoyable. I also greatly appreciated the support of my supervisor from the DORBINE project, Mario, who was always available to answer my questions and help me work through any challenges. The internship allowed me to apply my existing knowledge in a practical environment while developing new technical skills. Overall, I am very grateful for the opportunity, the guidance I received, and the valuable experience I gained during my time there.”

Gabriel Puri: Over the past four weeks as an intern, I’ve had a wonderful experience. I had the opportunity to work on streaming immersive media to the Apple Vision Pro, and I even created a Swift application that streams spatial videos to a local server. This server processes the incoming stream using my integrated pipeline, which enables adaptive bitrate streaming. I’ve learned a lot about encoding and how it’s done in real-world applications such as Netflix streams. I’ve also learned through trial and error, for example by trying an approach and failing, but eventually succeeding. I also learned how to accurately grade video quality via AVQT. One of the most memorable aspects of the internship was undoubtedly the incredibly welcoming and inclusive team, as well as my supervisor and mentor, Kamran from the ATHENA project. He is a true expert in his field and provided me with crucial support. This internship has allowed me to apply my interest in computer science to useful real-world scenarios and gain industry insights. Overall, I am incredibly grateful for this opportunity and for all the guidance and support I received throughout the internship.

 

Posted in ATHENA | Comments Off on Interns at ATHENA (Summer 2026)

Self-Training for Content-Aware Video Quality Enhancement in HTTP Adaptive Streaming

Self-Training for Content-Aware Video Quality
Enhancement in HTTP Adaptive Streaming

IEEE Transactions on Broadcasting

[PDF]

Yiying Wei (AAU, Austria), Hadi Amirpour (AAU, Austria),  Wei Zhou (Cardiff University, UK), Wassim Hamidouche (TII, UAE) and Christian Timmerer (AAU, Austria)

Abstract: Fluctuations in video segment download rates and resolution switching in HTTP Adaptive Streaming (HAS) make it challenging to maintain a consistent Quality of Experience (QoE). However, the impact of such switching is often underestimated, and broadly applicable mitigation strategies remain underexplored. In the past, content-aware approaches have been introduced, using deep neural networks (DNNs) trained on an individual video segment to enhance its quality. These DNNs, transferred as a model stream alongside the video bitstream, allow clients to improve playback quality. However, transferring model streams adds bitrate overhead and additional architectural components, limiting practical use. Furthermore, supporting a wide range of device capabilities with a single DNN is impractical, as it would require device-specific models for each configuration, an approach that becomes unmanageable with increasing device heterogeneity. In this paper, we propose a new self-training method that enables clients to train content-aware video super-resolution (SR) models locally by leveraging previously downloaded high-quality segments. These segments are downscaled and used to train lightweight DNNs, which are then applied to enhance subsequent lower-quality segments. To keep training efficient and real-time, we select only a few predefined frames and extract the most informative patches using a lightweight sampling strategy. Experiments demonstrate that this approach significantly improves visual quality, with average PSNR gains of 1.07 dB (2× upscaling), 0.43 dB (3×), and 0.58 dB (4×) using ESPCN, a lightweight SR approach. To further validate the effectiveness of our approach, we conducted a series of ablation studies to analyze the contributions of individual components. Real-device measurements and end-to-end HAS simulations further show that self-training requires only 2.9–13.0% of a 4-second segment interval on CPU and 1.1–5.3% on GPU across tested mobile devices, while improving VMAF/QoE with only marginal additional rebuffering compared with generic SR.

Posted in ATHENA | Comments Off on Self-Training for Content-Aware Video Quality Enhancement in HTTP Adaptive Streaming

DRLM: Deep Reinforcement Learning-Based LLM Query Orchestration in Edge Environments

DRLM: Deep Reinforcement Learning-Based LLM Query Orchestration in Edge Environments

Reza Farahani (TU Wien, Austria), Zoha Azimi (AAU, Austria), Mario Colosi (University of Messina, Italy), Lauri Lovén (University of Oulu, Finland), Christian Timmerer (AAU, Austria), Schahram Dustdar (TU Wien, Austria)

IEEE Global Communications Conference (GLOBECOM)

7 – 11 December 2026

Macau S.A.R., China

Large language model (LLM) services increasingly process heterogeneous queries with diverse latency, accuracy, and resource requirements. While edge deployment reduces response time, the heterogeneity of devices and the diversity of model families, parameter scales, and quantization levels make efficient LLM query orchestration challenging. This paper introduces DRLM, a Deep Reinforcement Learning-based LLM query orchestration framework in edge environments. DRLM integrates two lightweight predictors: (i) a class-conditioned quality estimator that maps queries to semantic categories and infers model performance, and (ii) a feature-driven latency predictor that estimates inference time across model-device configurations. These predictions, combined with system state (resource utilization and queue dynamics), feed a factorized Proximal Policy Optimization (PPO) agent that performs state-aware orchestration decisions. To enable data-driven orchestration, we construct a large-scale benchmarking dataset with 223 835 measurements spanning 1258 queries, 6 query classes, 8 model families (32 deployed instances), 5 quantization levels, and heterogeneous edge devices. Evaluation on a realistic 64-node edge cluster and comparison with three baselines and two state-of-the-art methods show that DRLM reduces inference latency by up to 51 % and queuing delay by up to 67 %, while incurring at most 8 % accuracy loss. DRLM further improves latency under increasing workloads up to 61.4 %, demonstrating robust and stable orchestration.

Posted in ATHENA | Comments Off on DRLM: Deep Reinforcement Learning-Based LLM Query Orchestration in Edge Environments

VCA-FFMPEG: A FFmpeg Filter for Video Complexity Analysis

VCA-FFMPEG: A FFmpeg Filter for Video Complexity Analysis

34th ACM International Conference on Multimedia 2026 (ACM MM 2026)

10–14 November 2026

Rio de Janeiro, Brazil

[PDF]

 Hadi Amirpour (AAU, Austria) Mykyta Skipenko (AAU, Austria), and Christian Timmerer (AAU, Austria)

Abstract:

VCA is a widely used open-source framework for estimating spatial and temporal video complexity features for applications such as per-title encoding, bitrate ladder generation, and content-adaptive streaming. However, the original VCA workflow requires videos to be preprocessed into raw formats such as YUV or Y4M before analysis, introducing additional decoding, conversion, storage, and pipeline overhead. Moreover, the standalone implementation has limited portability across platforms, devices, and heterogeneous multimedia processing systems.

In this paper, we present VCA-FFmpeg, an open-source integration of VCA as a native filter inside the FFmpeg multimedia framework. By embedding VCA directly into FFmpeg’s filtering pipeline, the proposed system removes the need for external preprocessing and intermediate format conversion, enabling complexity analysis on virtually any video format supported by FFmpeg. The integration allows users to extract spatial, temporal, brightness, and chroma descriptors during standard decoding, transcoding, or filtering operations using a simple command-line interface.

VCA-FFmpeg supports configurable block-based analysis, multi-threaded execution, optional low-pass DCT acceleration, SIMD optimizations, and YUView-compatible block-level visualization. Since FFmpeg is broadly supported across operating systems and hardware platforms, the proposed framework improves portability, usability, and reproducibility compared to standalone VCA. In our runtime evaluation, VCA-FFmpeg processed MP4 input at 462.7 fps, compared with 194.5 fps for the standalone workflow that first converts MP4 to Y4M and then applies VCA. By leveraging FFmpeg’s multimedia ecosystem, VCA-FFmpeg provides an efficient and scalable solution for in-pipeline video complexity feature extraction in multimedia systems, adaptive streaming, and video coding research.

Posted in ATHENA | Comments Off on VCA-FFMPEG: A FFmpeg Filter for Video Complexity Analysis