Configurable Hierarchical Wafer scale Switch Architecture for Scalable High-Bandwidth Interconnect Systems
Main Article Content
Abstract
This paper provided a hierarchical wafer scale switching architecture which is configurable and scalable and allows communication over a large number of dies at high bandwidth. The structure of the proposed fabric is legislation into the form of Outside, Middle, and Inside Switch Clusters connected with the help of special inter-cluster wiring modules. The architecture, unlike a classic flat two-dimensional mesh, separates the boundary traffic processing, middle redistribution and inner packet forwarding into two layers of communication. Every switch die has three programmable states of the ports unused, packet-switched, and circuit-switched. Packet-switched packet is handled by destination extraction, route-table shared, arbitration and crossbar forwarding, whereas circuit-switched packet take a low-overhead fixed route. Port mode, circuit destination and route-table entries are programmed by a global configuration controller. A shared, dual-table routing structure is used at die level in order to minimize hardware duplication and be available to all active input ports. The RTL-based modular design offers a platform that is sensible to assess hierarchical routing, customizable communication, and a huge scale switchover. A description of the system architecture, packet flow, routing strategy, arbitration mechanism, implementation methodology and evaluation metrics are provided in the paper. The results of simulations are deliberately intended to be inserted only after verification of simulation and synthesis.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
References
T. Li, J. Hou, J. Yan, R. Liu, H. Yang, and Z. Sun, “Chiplet heterogeneous integration technology—Status and challenges,” Electronics, vol. 9, no. 4, p. 670, 2020.
S. Y. Hou et al., “Wafer-level integration of an advanced logic-memory system through the second-generation CoWoS technology,” IEEE Transactions on Electron Devices, vol. 64, no. 10, pp. 4071–4077, 2017.
S. -R. Chun et al., “InFO_SoW (system-on-wafer) for high performance computing,” in Proc. IEEE ECTC, 2020, pp. 1–6.
R. Mahajan et al., “Embedded multi-die interconnect bridge (EMIB)—A high density, high bandwidth packaging interconnect,” in Proc. IEEE ECTC, 2016, pp. 557–565.
S. Pal et al., “Architecting waferscale processors—A GPU case study,” in Proc. IEEE HPCA, 2019, pp. 250–263.
W. Wang et al., “Demonstration of a wafer-level integration for system-on-wafer architecture,” in Proc. ICEPT, 2023, pp. 1–4.
S. Lie, “Cerebras architecture deep dive: First look inside the HW/SW co-design for deep learning,” IEEE Hot Chips 34, 2022.
B. Chang, R. Kurian, D. Williams, and E. Quinnell, “DOJO: Supercompute system scaling for ML training,” IEEE Hot Chips 34, 2022.
Y. Feng and K. Ma, “Chiplet actuary: A quantitative cost model and multi-chiplet architecture exploration,” in Proc. ACM/IEEE DAC, 2022, pp. 121–126.
S. Pal et al., “Designing a 2048-chiplet, 14336-core waferscale processor,” in Proc. ACM/IEEE DAC, 2021,
pp. 1183–1188.
J. Flich et al., “A survey and evaluation of topology-agnostic deterministic routing algorithms,” IEEE Transactions on Parallel and Distributed Systems, vol. 23, no. 3, pp. 405–425, 2012.
J. Chen, C. Li, and P. Gillard, “Network-on-chip topologies and performance: A review,” in Proc. NECEC, 2011, pp. 1–6.
L. Benini and G. De Micheli, “Networks on chips: A new SoC paradigm,” Computer, vol. 35, no. 1, pp. 70–78, 2002.
C. E. Leiserson, “Fat-trees: Universal networks for hardware-efficient supercomputing,” IEEE Transactions on Computers, vol. C-34, no. 10, pp. 892–901, 1985.
W. J. Dally and B. Towles, Principles and Practices of Interconnection Networks. Elsevier, 2003.