In traditional vision systems, the optical information is captured by a frame-based digital camera, and then the digital signal is processed afterwards using machine-learning algorithms. In this scenario, a large amount of data (mostly redundant) has to be transferred from a standalone sensing elements to the processing units, which leads to high latency and power consumption.