Skip to main navigation Skip to search Skip to main content

Intra coding complexity reduction in high efficiency video coding using RDO cost modeling and deep reinforcement learning

  • Mohammadreza Jamali

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Video compression technology has gained much attention in recent years due to the everincreasing popularity of high-definition (HD) and ultra-HD video applications and increased processing power of hardware and software. High efficiency video coding (HEVC)/H.265 is the most recent video coding standard which achieves a significant improvement in compression efficiency as compared to the other standards and provides a 50% bit rate reduction compared to the well-known H.264/advanced video coding (AVC) with the same quality. The performance improvement of HEVC is at the expense of much higher computational complexity at the encoder, making it challenging to deploy HEVC in real-time applications. In particular, HEVC increases the number of intra coding modes to 35, providing higher coding efficiency than the other video coding standards while increasing encoder complexity, which is mostly due to mode decision process by highly resource-demanding rate-distortion optimization (RDO). In addition, in frame splitting process, H.264/AVC employs 16 × 16 macroblocks, while HEVC introduces coding tree units (CTUs) with a maximum size of 64×64. The CTU may be split recursively and content-adaptively into coding units (CUs) in a quadtree-based manner, resulting in an efficient coding of background regions and objects with various sizes and shapes. In addition to the complexity imposed by mode decision, the frame partitioning process results in a significant computational complexity. In view of this, in this thesis, the HEVC intra coding is studied and multiple novel methods are proposed to reduce its computational complexity and encoding time. The proposed methods are revolving around two areas of mode decision and CU size decision. The first proposed method is based on the prediction of the RDO cost by a low-complexity SATD-based metric. Through predicting the RDO cost, the non-promising modes are discarded from further processing giving rise to substantial computation saving. This method provides a 30% time reduction with a 0.8% Bjøntegaard delta rate (BD-Rate) increase as compared to HEVC test model (HM); leading to a desirable trade-off. In the second contribution, a mode classification in chroma coding is proposed to adaptively reduce chroma intra modes based on block texture. As compared to HM, the chroma mode decision method provides a 6% time reduction with a 0.07% BD-Rate increase. The third contribution, in this thesis, is a gradient-based method, using the Prewitt operator, to eliminate the non-relevant directional modes from the list of candidates. The proposed method achieves a time reduction of 11.4% with a BD-Rate increase of 0.62% in comparison to HM. In the fourth proposed method, the most relevant modes of the neighboring blocks are considered to exploit the spatial redundancy across a frame. A classification of SATD costs is also proposed which permits the elimination of several candidate modes prior to RDO. It is shown that these two approaches, combined with the gradient-based algorithm, provide a 35.6% time reduction with a 1.07% BD-Rate loss. The proposed mode decision methods are combined, resulting in a 47.3% encoding time reduction with a quality loss of 1.37% BD-Rate as compared to the HM. The fifth contribution, in this thesis, is a fast intra coding method based on global and directional gradients to early terminate the CU splitting and avoid performing the highcomplexity RDO process for the next CU levels. This approach, combined with mode decision, reduces the encoding time by 52% on average, with a small quality loss of 1.50% BD-Rate. In the sixth contribution, a method based on the Bayesian classification is proposed to reduce the complexity of CU splitting process. Two binary classification problems are considered for early splitting and early splitting termination. It is shown that using the proposed method a 43.2% time reduction with a quality loss of 1.07% BD-Rate can be achieved. The seventh contribution is a CU size decision method based on reinforcement learning, active feature acquisition and neural networks. This method carries out early splitting and early splitting termination by considering the encoder and CU as an agent-environment system. The proposed method provides a 51.3% time reduction and a 0.84% BD-Rate loss. In addition, combining the proposed mode decision methods with this novel approach gives a total time reduction of 62.4% and a BD-Rate loss of 1.23% comparing to HM.
Date30 Aug 2018
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorStéphane Coulombe (Supervisor)

Cite this

'