Skip to main navigation Skip to search Skip to main content

Conception d’une mémoire cache L1 asynchrone pour un processeur ARM

Translated title of the thesis: Design of an asynchronous L1 cache for an ARM processor
  • Louis-Charles Trudeau

Student thesis: Master's thesisMaster in Engineering: Electrical Engineering

Abstract

As part of a global research program aimed at adapting Octasic’s asynchronous architecture to the implementation of general purpose processors, this work targets the design of a level one (L1) cache for ARM-like processors. The overall goal is to obtain processors that are efficient in terms of energy and size while remaining competitive in terms of computing power. Octasic current processors separate the asynchronous CPU core from the synchronous L1 instruction and data caches. The CPU core is segmented in multiple execution units (XUs) that are synchronized by token rings that are based on transition signaling. A token is assigned to each resource (Instruction Fetch, Register Read/Write, Jump/Conditionnal Branch, Data Memory, etc.) and shared among the XUs. Tokens are used and released according to a variable delay that matches the resource’s combinational logic. This asynchronous architecture reduces energy consumption by half for the same computing power when compared to its synchronous counterparts. This work focuses on improving the memory access. Currently, the asynchronous CPU core accesses a synchronous L1 cache. Fetching an instruction or accessing data imposes a 2-cycle synchronization penalty, thus resulting in a performance degradation. The primary objective of this work is to mitigate this latency. Also, the synchronous cache clock network still accounts for a substantial portion of the processor’s energy consumption, even though efficient clockgating strategies are used. The second objective is to increase the cache energy efficiency by removing the substantial clock tree while reducing the cache pipeline complexity. Furthermore, another objective is to increase the data transfer speed between the L1 and L2 memory by eliminating the need to share a global clock network. This thesis introduces a new self-timed pipeline developed for an L1 instruction cache. The self-timed pipeline integrates features from Octasic’s token-based architecture and Click elements. Simulation and post-layout timing were performed using a 28nm bulk process. Analysis have shown that energy efficiency can be improved by over 21% while reducing memory access time by more than 25% in average. The gathered results have also shown that the designed cache is specifically efficient when the processor’s execution pipeline flush is taken into account, thus far reducing dynamic power consumption by 68%. The asynchronous L1 cache designed during this Master’s degree has been published at the 21st IEEE International Symposium on Asynchronous Circuits and Systems (IEEE ASYNC 2015).
Date15 Dec 2015
Original languageFrench
Awarding Institution
  • École de technologie supérieure
SupervisorFrançois Gagnon (Supervisor) & Ghyslain Gagnon (Co-supervisor)

Cite this

'