ICIP 2024 Grand Challenge on Video Complexity

IEEE International Conference on Image Processing (IEEE ICIP)

Grand Challenge on

Video Complexity

27-30 October 2024, Abu Dhabi, UAE

https://cd-athena.github.io/GCVC

 

Organizers:

  • Ioannis Katsavounidis (Meta, USA)
  • Hadi Amirpour (AAU, Austria)
  • Ali Ak (Nantes Univ., France)
  • Anil Kokaram (TCD, Ireland)
  • Christian Timmerer (AAU, Austria)

 

Abstract: Video compression standards rely heavily on eliminating spatial and temporal redundancy within and across video frames. Intra-frame encoding targets redundancy within blocks of a single video frame, whereas inter-frame coding focuses on removing redundancy between the current frame and its reference frames. The level of spatial and temporal redundancy, or complexity, is a crucial factor in video compression. Generally, videos with higher complexity require a greater bitrate to maintain a specific quality level. Understanding the complexity of a video beforehand can significantly enhance the optimization of video coding and streaming workflows. While Spatial Information (SI) and Temporal Information (TI) are traditionally used to represent video complexity, they often exhibit low correlation with actual video coding performance. In this challenge, the goal is to find innovative methods that can quickly and accurately predict the spatial and temporal complexity of a video, with a high correlation to actual performance. These methods should be efficient enough to be applicable in live video streaming scenarios, ensuring real-time adaptability and optimization.

 

Posted in ATHENA | Comments Off on ICIP 2024 Grand Challenge on Video Complexity

Beyond Curves and Thresholds – Introducing Uncertainty Estimation to Satisfied User Ratios for Compressed Video

Picture Coding Symposium (PCS) 

12-14 June 2024, Taichung, Taiwan

[PDF]

Jingwen Zhu (University of Nantes, France), Hadi Amirpour (AAU, Austria), Raimund Schatz (AIT, Austria), Patrick Le Callet (University of Nantes, France)and Christian Timmerer (AAU, Austria)

Abstract: Just Noticeable Difference (JND) establishes the threshold between two images or videos wherein differences in quality remain imperceptible to an individual. This threshold, collectively known as the Satisfied User Ratio (SUR), holds significant importance in image and video compression applications, ensuring that differences in quality are imperceptible to the majority (p%) of users, known as p%SUR. While substantial efforts have been dedicated to predicting the p%SUR for various encoding parameters (e.g., QP) and quality metrics (e.g., VMAF), referred to as proxies, systematic consideration of the prediction uncertainties associated with these proxies has hitherto remained unexplored. In this paper, we analyze the uncertainty of p%SUR through Confidence Interval (CI) estimation and assess the consistency of various Video Quality Metrics (VQMs) as proxies for SUR. The analysis reveals challenges in directly using p%SUR as ground truth for training models and highlights the need for uncertainty estimation for SUR with different proxies.

Posted in ATHENA | Comments Off on Beyond Curves and Thresholds – Introducing Uncertainty Estimation to Satisfied User Ratios for Compressed Video

ICIP 2024 Grand Challenge on 360° Video Super Resolution

IEEE International Conference on Image Processing (IEEE ICIP)

Grand Challenge on

360° Video Super Resolution and Quality Enhancement

27-30 October 2024, Abi Dhabi, UAE

https://www.icip24-video360sr.ae/home

 

Abstract: Omnidirectional visual content, commonly referred to as 360-degree images and videos, has garnered significant interest in both academia and industry, establishing itself as the primary media modality for VR/XR applications. 360-degree videos offer numerous features and advantages, allowing users to view scenes in all directions, providing an immersive quality of experience with up to 3 degrees of freedom (3DoF). When integrated on embedded devices with remote control, 360-degree videos offer additional degrees of freedom, enabling movement within the space (6DoF). However, 360-degree videos come with specific requirements, such as high-resolution content with up to 16K video resolution to ensure a high-quality representation of the scene. Moreover, limited bandwidth in wireless communication, especially under mobility conditions, imposes strict constraints on the available throughput to prevent packet loss and maintain low end-to-end latency. Adaptive resolution and efficient compression of 360-degree video content can address these challenges by adapting to the available throughput while maintaining high video quality at the decoder. Nevertheless, the downscaling and coding of the original content before transmission introduces visible distortions and loss of information details that cannot be recovered at the decoder side. In this context, machine learning techniques have demonstrated outstanding performance in alleviating coding artifacts and recovering lost details, particularly for 2D video. Compared to 2D video, 360-degree video presents a lower angular resolution issue, requiring augmentation of both the resolution and the quality of the video. This challenge presents an opportunity for the scientific research and industrial community to propose solutions for quality enhancement and super-resolution for 360-degree videos.

 

Posted in ATHENA | Comments Off on ICIP 2024 Grand Challenge on 360° Video Super Resolution

Prof. Mohammad Ghanbari (1948-2024)

In the wake of the passing of Prof. Mohammad Ghanbari, we extend our deepest condolences to his family during this challenging time. Prof. Ghanbari was a distinguished member of our Christian Doppler Laboratory ATHENA since its inception, and we consider ourselves privileged to have had the opportunity to collaborate with him. His contributions, reflected in over 30 joint publications in video coding and streaming, have been accepted at renowned publication venues such as IEEE TCSVT, ACM TOMM, IEEE TIP, IEEE TNSM, IEEE ICIP, IEEE ICASSP, IEEE ICME, ACM MMSys, PCS, IEEE MMSP, among others. Prof. Ghanbari played a pivotal role in the success of our research endeavors, and his profound knowledge, insightful input, and invaluable guidance were consistently valued.

The entire Institute for Information Technology, especially those at the Christian Doppler Laboratory ATHENA, feels deeply saddened by the loss of Prof. Ghanbari. As we come to terms with this period of mourning, reflection, and farewell, we extend our warmest wishes and heartfelt sympathies to his family and the wider research community.

Posted in ATHENA | Comments Off on Prof. Mohammad Ghanbari (1948-2024)

DIGITWISE: Digital Twin-based Modeling of Adaptive Video Streaming Engagement

The 15th ACM Multimedia Systems Conference

15-18 April, 2024 | Bari, Italy

Conference website

[PDF]

Emanuele Artioli (AAU, Austria), Farzad Tashtarian (AAU, Austria), and Christian Timmerer (AAU, Austria)

Abstract:

As the popularity of video streaming entertainment continues to grow, understanding how users engage with the content and react to its changes becomes a critical success factor for every stakeholder. User engagement, i.e., the percentage of video the user watches before quitting, is central to customer loyalty, content personalization, ad relevance, and A/B testing. This paper presents DIGITWISE, a digital twin-based approach for modeling adaptive video streaming engagement. Traditional adaptive bitrate (ABR) algorithms assume that all users react similarly to video streaming artifacts and network issues, neglecting individual user sensitivities. DIGITWISE leverages the concept of a digital twin, a digital replica of a physical entity, to model user engagement based on past viewing sessions. The digital twin receives input about streaming events and utilizes supervised machine learning to predict user engagement for a given session. The system model consists of a data processing pipeline, machine learning models acting as digital twins, and a unified model to predict engagement. DIGITWISE employs the XGBoost model in both digital twins and unified models. The proposed architecture demonstrates the importance of personal user sensitivities, reducing user engagement prediction error by up to 5.8% compared to non-user-aware models. Furthermore, DIGITWISE can optimize content provisioning and delivery by identifying the features that maximize engagement, providing an average engagement increase of up to 8.6 %.

Keywords: digital twin, user engagement, xgboost

Posted in ATHENA | Comments Off on DIGITWISE: Digital Twin-based Modeling of Adaptive Video Streaming Engagement

E-WISH: An Energy-aware ABR Algorithm For Green HTTP Adaptive Video Streaming

ACM Mile-High Video 2024

February 11-14, 2024, Marriott DTC, Denver, US

[PDF]

Daniele Lorenzi (AAU, Austria), Minh Nguyen (AAU, Austria), Farzad Tashtarian (AAU, Austria), and Christian Timmerer (AAU, Austria)

Abstract:

HTTP Adaptive Streaming (HAS) is the de-facto solution for delivering video content over the Internet. The climate crisis has highlighted the environmental impact of information and communication technologies (ICT) solutions and the need for green solutions to reduce ICT’s carbon footprint. As video streaming dominates Internet traffic, research in this direction is vital now more than ever. HAS relies on Adaptive BitRate (ABR) algorithms, which dynamically choose suitable video representations to accommodate device characteristics and network conditions. ABR algorithms typically prioritize video quality, ignoring the energy impact of their decisions. Consequently, they often select the video representation with the highest bitrate under good network conditions, thereby increasing energy consumption. This is problematic, especially for energy-limited devices, because it affects the device’s battery life and the user experience. To address the aforementioned issues, we propose E-WISH, a novel energy-aware ABR algorithm, which extends the already-existing WISH algorithm to consider energy consumption while selecting the quality for the next video segment. According to the experimental findings, E-WISH shows the ability to improve Quality of Experience (QoE) by up to 52% according to the ITU-T P.1203 model (mode 0) while simultaneously reducing energy consumption by up to 12% with respect to state-of-the-art approaches.

Keywords: HTTP adaptive streaming, Energy, Adaptive Bitrate (ABR), DASH

Posted in ATHENA | Comments Off on E-WISH: An Energy-aware ABR Algorithm For Green HTTP Adaptive Video Streaming

Optimal Quality and Efficiency in Adaptive Live Streaming with JND-Aware Low Latency Encoding

MHV 2024: ACM Mile High Video

11 – 14 Feb 2024 |  Denver, United States

Conference Website

[PDF][Slides]

Vignesh V Menon (Fraunhofer HHI),  Jingwen Zhu (École Centrale Nantes), Prajit T Rajendran (Université Paris-Saclay),   Samira Afzal (Alpen-Adria-Universität Klagenfurt), Klaus Schoeffmann (Alpen-Adria-Universität Klagenfurt), Patrick Le Callet (École Centrale Nantes), and Christian Timmerer (Alpen-Adria-Universität Klagenfurt)

Abstract: In HTTP adaptive live streaming applications, video segments are encoded at a fixed set of bitrate-resolution pairs known as bitrate ladder. Live encoders use the fastest available encoding configuration, referred to as preset, to ensure the minimum possible latency in video encoding. However, an optimized preset and optimized number of CPU threads for each encoding instance may result in (i) increased quality and (ii) efficient CPU utilization while encoding. For low latency live encoders, the encoding speed is expected to be more than or equal to the video framerate. To this light, this paper introduces a Just Noticeable Difference (JND)-Aware Low latency Encoding Scheme (JALE), which uses random forest-based models to jointly determine the optimized encoder preset and thread count for each representation, based on video complexity features, the target encoding speed, the total number of available CPU threads, and the target encoder. Experimental results show that, on average, JALE yield a quality improvement of 1.32 dB PSNR and 5.38 VMAF points with the same bitrate, compared to the fastest preset encoding of the HTTP Live Streaming (HLS) bitrate ladder using x265 HEVC open-source encoder with eight CPU threads used for each representation. These enhancements are achieved while maintaining the desired encoding speed. Furthermore, on average, JALE results in an overall storage reduction of 72.70%, a reduction in the total number of CPU threads used by 63.83%, and a 37.87% reduction in the overall encoding time, considering a JND of six VMAF points.

Keywords: Live streaming, low latency, encoder preset, CPU threads, HEVC.

Posted in GAIA | Comments Off on Optimal Quality and Efficiency in Adaptive Live Streaming with JND-Aware Low Latency Encoding