Quantization and Fast Inference: A Practitioner's Guide to Efficient AI, (Paperback)

★★★★☆ 4.0 41 reviews

US$24.00
Price when purchased online
Free shipping Free 30-day returns

Sold and shipped by theologieducorps.com
We aim to show you accurate product information. Manufacturers, suppliers and others provide what you see here.
US$24.00
Price when purchased online
Free shipping Free 30-day returns

How do you want your item?
You get 30 days free! Choose a plan at checkout.
Shipping
Arrives Jul 30
Free
Pickup
Check nearby
Delivery
Not available

Sold and shipped by theologieducorps.com
Free 30-day returns Details

Product details

Management number 238570852 Release Date 2026/07/11 List Price US$24.00 Model Number 238570852
Category

<b>Get the eBook free when you register your print book at Manning.</b> <p>Today's AI models demand a lot of memory, compute, and server horsepower--which quickly translates into cost. This book show you how you can optimize AI models without architectural redesigns or task-specific compression. It reveals practical techniques for quantization, systematically reducing numerical precision to achieve faster inference, lower memory usage, and cheaper deployment--all with minimal accuracy loss. </p><p>From quantization fundamentals to runtime packaging, the book gives you a complete and comprehensive overview of the full quantization pipeline. It starts by deriving quantization mapping from first principles, and then builds your knowledge and skill through techniques for production-tested PTQ and QAT workflows and a fully compressed deployment. You'll learn to apply post-training quantization to production models, run quantization-aware training using fake quantization and straight-through estimators, and handle subtle tradeoffs like activation outliers in LLMs, KV cache pressure, and sub-8-bit formats like NF4 and FP4. </p><p> <b>What's inside</b> </p><p> - Applying post-training quantization to production models<br> - Deploying efficiently on CPUs, edge devices, and mobile<br> - Framework-agnostic techniques and real cross-framework parity testing<br> - Flowcharts and checklists for efficient decision making </p><p><b>About the reader</b> </p><p> For ML engineers and researchers experienced in Python. </p><p> <b>About the author</b> </p><p> <b>Vivek Kalyanarangan</b> is an AI/ML architect, researcher, and educator with over twelve years of experience designing and deploying large-scale machine learning systems.</p>

  • Quantization and Fast Inference: A Practitioner's Guide to Efficient AI, (Paperback)
  • Author: Manning Publications
  • ISBN: 9781633433915
  • Format: Paperback
  • Publication Date: 2026-12-29
  • Page Count: 350
Book format Paperback
Fiction/nonfiction Non-Fiction
Genre Computing & Internet
Publication date December, 2026
Pages 350
Subgenre Data Science
Series title No Series
Number in series 0
Edition 1
Publisher Manning Publications
Language English
Is collectible N
Recording time 0 min
Retail packaging Single Piece
Assembled product dimensions (l x w x h) 7.38 x 6.00 x 9.25 in
Assembled product weight 0.92 lb
Bisac subject heading Computers

Correction of product information

If you notice any omissions or errors in the product information on this page, please use the correction request form below.

Correction Request Form

Customer ratings & reviews

4 out of 5
★★★★☆
41 ratings | 17 reviews
How item rating is calculated
View all reviews
5 stars
75% (31)
4 stars
8% (3)
3 stars
4% (2)
2 stars
2% (1)
1 star
11% (5)
Sort by

There are currently no written reviews for this product.