Cantitate/Preț
Produs

Self-restructuring in Fault Tolerant Architecture

Autor Itsuo Takanami
en Limba Engleză Paperback – 17 apr 2025

Adresat arhitecților de sistem și inginerilor hardware avansați, Self-restructuring in Fault Tolerant Architecture de Itsuo Takanami explorează o nișă critică în designul sistemelor de calcul paralele: capacitatea de auto-reparare la nivel de circuit. Într-o eră în care tehnologiile VLSI și WSI (Wafer Scale Integration) permit integrarea a sute de elemente de procesare pe un singur cip, fiabilitatea devine o provocare majoră. Notăm cu interes faptul că autorul mută paradigma de la intervenția software externă către mecanisme de control integrate direct în hardware.

Abordarea este una riguros tehnică, concentrându-se pe rețelele de procesatoare conectate în formă de matrice (mesh). Putem afirma că elementul distinctiv al lucrării rezidă în descrierea detaliată a algoritmilor de reconfigurare implementați prin circuite digitale, care permit izolarea automată a componentelor defecte și înlocuirea lor cu unități de rezervă (spares). Această metodă reduce drastic timpul de inactivitate al sistemului, fiind esențială în medii unde monitorizarea externă sau mentenanța manuală sunt imposibile.

Dacă Software Design for Resilient Computer Systems de Igor Schagaev v-a oferit cadrul teoretic asupra modului în care software-ul de sistem interacționează cu hardware-ul pentru toleranța la erori, lucrarea de față oferă instrumentele practice și specificațiile de proiectare pentru nivelul fizic. În timp ce Architecting Dependable Systems se concentrează pe arhitectura software generală, Itsuo Takanami analizează matematic probabilitatea de succes a reconfigurării în funcție de numărul de defecte și dispunerea fizică a rezervelor pe diagonală sau pe laturile matricei. Publicată de Springer, această lucrare de 124 de pagini reprezintă un ghid dens în specificații tehnice despre supraviețuirea sistemelor masiv paralele.

Citește tot Restrânge

Preț: 31225 lei

Preț vechi: 39032 lei
-20%

Puncte Express: 468

Carte disponibilă

Livrare economică 06-20 octombrie

Livrare prin curier în România Termenul estimat este afișat lângă disponibilitate.
Transport gratuit de la 40000 lei Plată online sau ramburs, în funcție de opțiunile comenzii.
Retur gratuit în 14 zile Comandă securizată și suport în română.

Specificații

ISBN-13: 9789819615384
ISBN-10: 9819615380
Pagini: 124
Dimensiuni: 155 x 235 x 7 mm
Greutate: 0.22 kg
Editura: Springer

De ce să citești această carte

Recomandăm această carte inginerilor hardware care proiectează sisteme critice unde fiabilitatea este vitală. Cititorul câștigă o înțelegere profundă a algoritmilor de auto-reconfigurare integrați, esențiali pentru tehnologiile VLSI moderne. Este o resursă tehnică valoroasă pentru implementarea sistemelor capabile să își mențină conectivitatea logică fără intervenție umană, optimizând consumul de energie și spațiul prin soluții de tip built-in circuit.


Descriere

Recently, high-speed and high-quality technologies for processing many kinds of information have become essential and will become more and more necessary in the future. For such needs, parallel computer systems composed of many processing elements (PEs) are used and it is important to make high reliable systems which is called "fault-tolerant computer systems". As VLSI technology has developed, the realization of parallel computer systems using multi-chip module (MCM) or wafer scale integration (WSI) has been considered so as to enhance the speed of the computers, decrease energy consumption and sizes, and so on. In such a realization, entire or significant parts of PEs and connections among them are connected or implemented on a board or wafer. Therefore, the reliability and/or yield of the system may become drastically low if there is no strategy for coping with faults or defects. In realizing such systems as well as parallel computer systems, in order to restore the correct computation capabilities of the systems with faults, it must be reconfigured appropriately using spare PEs so that the faulty PEs are eliminated from the computation paths by replacing faulty PEs with healthy spare PEs and the remaining healthy PEs maintain correct logical connectivity among them. Various strategies to reconfigure a faulty physical system into a fault-free target logical system are described in the literature. Some of these techniques employ very powerful reconfiguring systems that can repair a faulty processor array with almost certainty, even in the presence of clusters of multiple faults. However, the key limitation of these techniques is that they are executed in software programs to run on an external host computer and they cannot be designed and implemented efficiently within a system. If a faulty system can be self-reconfigured by a built-in circuit or network, the system down time is significantly reduced. Furthermore, the system will become more reliable when it is used in such environments that the fault information cannot be monitored externally and manual maintenance operations are difficult. This book concerns fault-tolerant systems consisting of many PEs, mainly mesh-connected processor arrays. A mesh-connected processor array is a kind of form of massively parallel computing systems which consist of hundreds of PEs and have regular and modular structures, small wiring length between PEs, and high scalabilities. Here, self-reconfiguration of processor systems with spares using built-in digital circuits are focused on and spare arrangements together with networks connecting among PEs and reconfiguration algorithms with digital circuits are described, considering the number of spares, reconfiguration algorithms and their hardware realizations (built-in circuits), etc. where spares are arranged on the sides or diagonal of arrays. The effectiveness of the systems is evaluated in terms of the survival rates (successfully reconfigured rates) for the number of faults, and the array reliabilities (successfully reconfigured probabilities under the condition that each PE is equally reliable).