PinnedInIntel Analytics SoftwarebyIntel(R) Neural Compressor & AutoRound·Oct 31, 2025Accelerating vLLM and SGLang Deployment using AutoRoundAutoRound: A Leading Quantization Algorithm for LLMs/VLMs
PinnedInIntel Analytics SoftwarebyIntel(R) Neural Compressor & AutoRound·Nov 1, 2022Personalized Stable Diffusion with Few-Shot Fine-TuningCreate Your Own Stable Diffusion on a Single CPUA response icon1A response icon1
Intel(R) Neural Compressor & AutoRound·Dec 28, 2025Diffusion in AutoRound: Low‑Bit Quantization for Video and Image Generation ModelsOverviewA response icon1A response icon1
InIntel Analytics SoftwarebyIntel(R) Neural Compressor & AutoRound·Nov 22, 202410 Tips for Quantizing LLMs and VLMs with AutoRoundAutoRound V0.4 has been released, featuring major updates to experimentally support Vision-Language Models (VLMs), and many quantized…
InIntel Analytics SoftwarebyIntel(R) Neural Compressor & AutoRound·Aug 16, 2024Quantization on Intel Gaudi Series AI AcceleratorsIntel Neural Compressor v3.0 Supports Quantization across Intel Hardware
InIntel Analytics SoftwarebyIntel(R) Neural Compressor & AutoRound·Jun 6, 2024Accelerating Qwen2 Models with Intel Extension for TransformersHigh Performance WOQ INT4 Inference on Intel Xeon Processors
InIntel Analytics SoftwarebyIntel(R) Neural Compressor & AutoRound·May 31, 2024Accelerating GGUF Models with TransformersImproving Performance and Memory Usage on Intel Platforms
InIntel Analytics SoftwarebyIntel(R) Neural Compressor & AutoRound·May 11, 2024Low-Bit Quantized Open LLM LeaderboardA New Tool to Find High-Quality Models for a Given Client
InIntel Analytics SoftwarebyIntel(R) Neural Compressor & AutoRound·Apr 2, 2024The AutoRound Quantization AlgorithmWeight-Only Quantization for LLMs Across Hardware PlatformsA response icon2A response icon2
InIntel Analytics SoftwarebyIntel(R) Neural Compressor & AutoRound·Mar 22, 2024Run LLMs on Intel GPUs Using llama.cppTaking Advantage of the New SYCL BackendA response icon1A response icon1