In recent years, the demand for high-quality video has increased and motivated video compression technologies to improve. High Efficiency Video Coding (HEVC) is the newest of such advances that intends to reduce the bit rate by half relative to H.264/MPEG-4 AVC for the same quality. This improvement comes with much higher computational complexity at the encoder; making it difficult to deploy HEVC in typical applications. Using heterogeneous architectures is a recognized approach to reduce the execution time of complex algorithms. However, HEVC is not well designed to be executed on massively parallel architectures. This research aims to study the fine-grained parallelization of rate-constrained motion estimation which is the most time-consuming part of the HEVC encoder.
In this project, we investigate the existing parallel tools in HEVC and the literature related to parallel implementations of HEVC. We discuss the drawbacks of existing methods. Then, we propose a two-stage parallel framework, which is flexible and efficient. The proposed framework provides a high degree of parallelism suitable for heterogeneous architectures. Furthermore, to reduce the rate-distortion (RD) performance loss caused by breaking data dependencies, we propose a multi-predictor rate-constrained motion estimation approach and a multiple temporal predictor method. According to the experimental results, our proposed methods improve the Bjøntegaard-Delta Rate (BD-Rate) by an average of 1.44% compared to the one predictor parallel rate-constrained motion estimation (RCME) method and 0.92% compared to a leading state-of-the-art method which uses the average of predictors. Moreover, according to the graphics processing unit (GPU) hardware specifications an innovative search method is introduced to exploit the power of GPUs more efficiently. The execution time of the whole encoding process is reduced 40% compared to the fastest RCME algorithm.
The results of this research are expected to lead to an improved architecture for HEVC encoders that can exploit the computational power of massively parallel many-core architectures to increase speed while preserving the RD performance.
| Date | 1 Jun 2018 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Stéphane Coulombe (Supervisor) & Carlos Vázquez (Co-supervisor) |
|---|
Hojati Najafabadi, E. (Author),
Coulombe (Supervisor) &
Vázquez (Co-supervisor),
1 Jun 2018Student thesis: Master's thesis › Master in Engineering: Engineering