HiPEAC

Hardware accelerators for low power deep neural networks in mixed signal and programmable SoCs

Since 2012, when a deep neural network (DNN) called AlexNet achieved much higher accuracy than any other previous algorithm tailored for image recognition, the paradigm of deep learning is monopolizing most of the research lines in the field of computer vision. Its basic principle of operation is the progressive automatic extraction of image features learnt during a prior process of training. In contrast, classical algorithms require an ad-hoc design of low-level features that constitute the input of shallow classifiers.

However, their implementation represents a great cost both in computational resources and power dissipation, especially in vision applications with speeds close to real time operation. The design of a network capable of inferring in real time is a real challenge, that becomes even greater if this network is implemented on low-power embedded devices. For these cases, intermediate solutions between massively parallel computers such as GPUs and chips implementing ad-hoc solutions must be explored. In this sense, this project focuses on the design of hardware accelerators for DNNS devoted to image processing tasks. Given the algorithmic unified approach of the DNNs, the proposed solutions can be used systematically at different levels.

However, our proposal does not focus on a block that is isolated from the rest of the vision system, but adopts a holistic approach to address the problem. Considering the vision system as a whole, and taking as initial point the analysis at various levels of its architecture and its data flows, we will propose accelerators that optimize both the processing elements and the memory resources as well as their interconnections. Given the fact that the greater computational load of a DL network relies on Multiply and Accumulation operations and that they require a large number of weights stored in memory, we find also that memory bandwidth often becomes the bottleneck of the network. The solution does not only refer to the redesign of the memory hierarchy, but also to the analysis of the dataflow at different levels to exploit the parallelism of all resources. For this reason, it is necessary to implement strategies that allow the redistribution of resources, such as the placement of processing close to the sensor, the use of optimized memory hierarchies or the integration of part of computing in memory. Moreover, the analog conversion-information at the focal plane level can be used to reduce the data flow and maximize the whole system parallelism. CMOS technology allows the integration in a same chip both sensing devices and converter and processing circuits. This represents a major challenge for electronic designers since it opens up a wide range of solutions regarding the imaging systems configuration.

In this project, we will take advantage of our experience in mixed signal microelectronic design and hardware accelerators for image processing to propose new hardware accelerator circuits and architectures for DNNs. These circuits could be integrated into low-power and real-time embedded vision systems. To demonstrate the feasibility of this proposal, a demonstrator of the image processing system for intelligent transport applications will be developed.


Expertise areas

Application areas: Transportation

Topics: Approximate computing, Energy efficiency / Low-power computing, FPGAs, Machine learning