Diffractive tensorized unit for million-TOPS general-purpose computing
This paper, published in Nature Photonics, was completed by a team led by Academician Dai Qionghai and Professor Fang Lu from Tsinghua University. It proposes and experimentally verifies a novel photonic processor architecture called the Diffraction Tensor Unit (DTU), aiming to solve the long-standing core challenges of non-reconfigurability and limited scalability in diffractive optical computing. For the first time, a high-performance diffractive photonic processor supporting general-purpose computing has been implemented on a chip. The core innovations can be divided into two parts: near-nuclear modulation and tensor processing.
Architectural Innovations: In the near-nuclear modulation problem, unlike directly controlling the millions of static, fixed diffractive neurons inside the diffraction core, a combination of Dynamic Diffraction Core (DDTC) and Static Diffraction Core (SDTC) is used to reconstruct the actual transmission matrix through local control. The figure shows the specific process of DTU reconstruction using DDTC.

The transmission process of a DTU can be interpreted as follows:
It includes signal input and parameter input. Let be the transmission system, matching the dimension of . This operation can be further refined as:
For , assuming the equivalent transfer matrix is , the main challenge in designing a DDTC is whether there exists a solution such that for any given , there exists at least one solution such that . That is, whether, through the design of a static system, with the assistance of dynamic parameters, for the input signal, the output, i.e., the transfer matrix, can be achieved .
By introducing , we can achieve , which means changing the function of the system and reconstructing the matrix. At this point:
Let , then:
The problem becomes whether a corresponding solution can be found such that this system of linear equations has a solution. A necessary and sufficient condition is that the number of parameter channels P is not less than the number of output channels O. And the rank of the matrix is equal to the number of channels O.
Based on the implementation of a single DTU, large-scale data processing can be achieved through the Zhang Lianghua architecture. The DTU uses a tensor decomposition method. The massive computational task is decomposed into multiple small tensor cores, which are then mapped to a computing cluster composed of multiple DTCs for parallel processing, giving the DTU excellent scalability.

Performance Verification Results:
Ø Basic General Computing Capabilities: DTU can perform matrix multiplication of any size 1024*1024 with a mean square error of 10⁻⁶ (based on 32 times time-domain multiplexing using 32*32 DTC).
Ø Advanced AI Task Implementation: DTU implements complex AI applications such as natural language generation, cross-modal recognition, image classification, and video generation. In the experiments, inference experiments were completed for MNIST and Fashion-MNIST image classification and natural language generation tasks, achieving accuracies of 97.7%, 85.4%, and 58.6%, respectively, which highly agree with the simulation results, demonstrating the feasibility of the architecture.

In the Natural Language Generation (NLG) task, figures (a) and (b) illustrate how to use DTUs for word prediction. The input word sequence is encoded as an optical signal, which is then cyclically computed and modulated in a DTU array to ultimately generate the predicted next word. The training convergence curve in figure (c) demonstrates that DTUs are comparable to Bi-LSTM/Transformers and have versatility. Figure (e) shows the evolution of word vectors with prediction. Figure (f) illustrates sentence generation. Figures (g) and (h) show the impact of the number of iterations and input length on accuracy. Figure (f) shows the results for Chinese poetry and couplets.

In the cross-modal recognition problem, video frames are serialized and tensed, and a cyclic DTU cluster is used to process the video content and generate text descriptions. The performance of the DTU is validated on standard datasets such as MSVD and MSR-VTT.

Two representative tasks, image classification and null logic generation (NLG), were implemented on a real-world DTU chip. On the MNIST handwritten digit classification task, the chip achieved an overall accuracy of 97.7%. On the NLG task, the actual output of the chip deviated from the simulation results by only 1.9%, and the average accuracy across a series of experiments reached 58.6%, very close to the 60.5% shown in the simulation results.
Original link:
Diffractive tensorized unit for million-TOPS general-purpose computing | Nature Photonics