Algorithms and Data Structures for Massive Datasets
Autor Dzejla Medjedovic, Emin Tahirovic Ilustrat de Ines Dedovicen Limba Engleză Paperback – 5 iul 2022
Ecosistemul analizei de date la scară largă necesită o schimbare de paradigmă față de algoritmii tradiționali, iar Algorithms and Data Structures for Massive Datasets oferă exact acest instrumentar tehnic. Notăm cu interes modul în care autorii, Dzejla Medjedovic și Emin Tahirovic, abordează limitările hardware atunci când seturile de date depășesc capacitatea RAM, concentrându-se pe structuri de date probabilistice și tehnici de „sketching”. Credem că structura progresivă a volumului facilitează înțelegerea unor concepte matematice complexe prin aplicabilitate practică. Prima parte explorează tehnici bazate pe hashing, precum Bloom Filters și HyperLogLog, esențiale pentru estimarea rapidă a apartenenței sau cardinalității. A doua parte se mută în zona fluxurilor de date (streaming), discutând eșantionarea și quantile-urile aproximative, în timp ce secțiunea finală este dedicată algoritmilor de memorie externă și structurilor de indexare precum LSM-Trees, fundamentale în bazele de date moderne de tip NoSQL. Cititorul care a aplicat ideile din Mining of Massive Datasets va găsi aici o completare tehnică riguroasă, axată mai puțin pe procesarea paralelă (MapReduce) și mai mult pe eficiența algoritmilor la nivel de structură de date individuală. De asemenea, spre deosebire de Small Summaries for Big Data, care oferă o introducere teoretică în sumarizarea datelor, volumul de față publicat de Manning Publications pune accent pe implementarea în sisteme reale, utilizând ilustrații și exemple din mediul de afaceri pentru a maximiza rata de transfer (throughput) a procesării.
Preț: 326.14 lei
Preț vechi: 407.67 lei
-20%
Carte disponibilă
Livrare economică 25 septembrie-09 octombrie
Specificații
ISBN-10: 1617298034
Pagini: 304
Dimensiuni: 191 x 237 x 20 mm
Greutate: 0.48 kg
Editura: Manning Publications
De ce să citești această carte
Recomandăm această carte inginerilor de date și dezvoltatorilor software care se confruntă cu limitări de memorie în procesarea volumelor mari de informații. Veți câștiga o înțelegere profundă a structurilor probabilistice care stau la baza sistemelor moderne de analiză, învățând cum să optimizați precizia în schimbul vitezei și al consumului redus de resurse. Este un ghid practic pentru construirea unor sisteme de date scalabile și eficiente.
Despre autor
Dzejla Medjedovic este profesor asociat, deținând un doctorat în informatică de la Universitatea Stony Brook, unde s-a specializat în algoritmi de memorie externă. Emin Tahirovic este, de asemenea, profesor și cercetător cu experiență în analiza datelor și bioinformatică, având un doctorat de la Universitatea din Pennsylvania. Expertiza lor combinată în algoritmi teoretici și aplicații practice se reflectă în claritatea cu care sunt explicate structurile de date complexe în acest volum, oferind cititorului o perspectivă academică ancorată în nevoile industriei software actuale.
Notă biografică
Cuprins
Descriere scurtă
In Algorithms and Data Structures for Massive Datasets you will learn:
Probabilistic sketching data structures for practical problems
Choosing the right database engine for your application
Evaluating and designing efficient on-disk data structures and algorithms
Understanding the algorithmic trade-offs involved in massive-scale systems
Deriving basic statistics from streaming data
Correctly sampling streaming data
Computing percentiles with limited space resources
Algorithms and Data Structures for Massive Datasets reveals a toolbox of new methods that are perfect for handling modern big data applications. You’ll explore the novel data structures and algorithms that underpin Google, Facebook, and other enterprise applications that work with truly massive amounts of data. These effective techniques can be applied to any discipline, from finance to text analysis. Graphics, illustrations, and hands-on industry examples make complex ideas practical to implement in your projects—and there’s no mathematical proofs to puzzle over. Work through this one-of-a-kind guide, and you’ll find the sweet spot of saving space without sacrificing your data’s accuracy.
About the technology
Standard algorithms and data structures may become slow—or fail altogether—when applied to large distributed datasets. Choosing algorithms designed for big data saves time, increases accuracy, and reduces processing cost. This unique book distills cutting-edge research papers into practical techniques for sketching, streaming, and organizing massive datasets on-disk and in the cloud.
About the book
Algorithms and Data Structures for Massive Datasets introduces processing and analytics techniques for large distributed data. Packed with industry stories and entertaining illustrations, this friendly guide makes even complex concepts easy to understand. You’ll explore real-world examples as you learn to map powerful algorithms like Bloom filters, Count-min sketch, HyperLogLog, and LSM-trees to your own use cases.
What's inside
Probabilistic sketching data structures
Choosing the right database engine
Designing efficient on-disk data structures and algorithms
Algorithmic tradeoffs in massive-scale systems
Computing percentiles with limited space resources
About the reader
Examples in Python, R, and pseudocode.
About the author
Dzejla Medjedovic earned her PhD in the Applied Algorithms Lab at Stony Brook University, New York. Emin Tahirovic earned his PhD in biostatistics from University of Pennsylvania. Illustrator Ines Dedovic earned her PhD at the Institute for Imaging and Computer Vision at RWTH Aachen University, Germany.
Table of Contents
1 Introduction
PART 1 HASH-BASED SKETCHES
2 Review of hash tables and modern hashing
3 Approximate membership: Bloom and quotient filters
4 Frequency estimation and count-min sketch
5 Cardinality estimation and HyperLogLog
PART 2 REAL-TIME ANALYTICS
6 Streaming data: Bringing everything together
7 Sampling from data streams
8 Approximate quantiles on data streams
PART 3 DATA STRUCTURES FOR DATABASES AND EXTERNAL MEMORY ALGORITHMS
9 Introducing the external memory model
10 Data structures for databases: B-trees, Bε-trees, and LSM-trees
11 External memory sorting