SendTech Times
News
MARKET SIGNAL:

Hugging Face Adds 4-Bit Nunchaku Loading To Diffusers

Newsroom brief

Hugging Face added Nunchaku Lite support to Diffusers, letting developers load 4-bit diffusion checkpoints with `from_pretrained()` while using Hub-delivered CUDA kernels.

Verified against source materialEdited by SendTech Times Chips & Compute DeskSource: Hugging Face
Hugging Face Adds 4-Bit Nunchaku Loading To Diffusers
Image source: Hugging Face

Text-to-image pipelines that can require 20GB to 30GB of VRAM now have a 4-bit loading path in Hugging Face Diffusers.

Hugging Face said in a July 23 post that the Nunchaku Lite integration is available through the standard `from_pretrained()` workflow.

The change targets model loading rather than only reducing stored weight files.

Kernels arrive through the Hub

Nunchaku Lite loads without a custom pipeline class, separate inference engine or local CUDA compilation, according to the post.

The runtime patches relevant `nn.Linear` modules in a stock Diffusers model with SVDQ or AWQ linear layers before loading the checkpoint, and CUDA kernels are downloaded through the `kernels` package.

The integration uses `svdq_w4a4` layers for attention and MLP projections, with INT4 and NVFP4 variants.

A second `awq_w4a16` path covers precision-sensitive normalization and modulation components in model families including FLUX and Qwen-Image.

The hardware table says NVFP4 checkpoints require NVIDIA Blackwell hardware, including RTX 50 series, RTX PRO 6000 and B200 GPUs, while INT4 variants support Turing, Ampere and Ada devices such as RTX 30 and 40 series cards, A100 and L40S.

Benchmarks stay vendor-owned

According to Hugging Face's example, one Nunchaku NVFP4 transformer paired with a bitsandbytes NF4 text encoder generated a square 1024-pixel image in roughly 1.7 seconds on an RTX 5090, while peak memory was about 12 GB; the same paragraph placed the BF16 pipeline at about 24 GB.

In a separate RTX PRO 6000 Blackwell benchmark, the blog listed a BF16 baseline at 3.00 seconds for the full pipeline and 31.1 GB peak VRAM.

The blog's benchmark table listed Nunchaku Lite NVFP4 at 2.27 seconds and 20.6 GB peak VRAM, Nunchaku Lite with `torch.compile` at 1.68 seconds and 20.6 GB, and Nunchaku Lite with an NF4 text encoder at 2.29 seconds and 16.0 GB.

The post characterized the result as up to 50% lower peak VRAM, roughly 30% better latency and as much as a 1.8x speedup with `torch.compile`; those figures remain vendor benchmark claims rather than independent lab measurements.

Quantization workflow widens

The post also points developers to `diffuse-compressor`, a companion toolkit for calibrating, quantizing, packaging and publishing new Diffusers models.

The FLUX.2 Klein 4B example workflow says the inspection stage should identify 100 SVDQ targets, 3 AWQ targets and 6 dense outer linears before quantization proceeds.

The remaining boundary is hardware and model structure.

The compatibility note excludes Volta and Hopper GPUs from current 4-bit kernel support, and the generic Nunchaku Lite path does not infer architecture-specific fused rewrites such as combined QKV projections.

That makes the integration useful as a standard Diffusers loading route, while specialized engines may still matter when a model family needs fused modules beyond the generic layer replacement path.

The post did not disclose production customer names for the new Diffusers path, so deployment proof is still limited to the published benchmark examples.

Share this article
inXf

Related articles

More
AMD Linux Patch Adds Low-Power CPU Core Support For Future Chips
Chips & Semiconductors

AMD Linux Patch Adds Low-Power CPU Core Support For Future Chips

AMD has submitted Linux kernel patches that add a low-power CPU core classification beside performance and efficiency cores, The public record still lacks the first processor, launch date or product line that will use the new core type.

Kioxia GP1 SSD Demo Pushes PCIe Gen6 Storage To 10 Million IOPS
Chips & Semiconductors

Kioxia GP1 SSD Demo Pushes PCIe Gen6 Storage To 10 Million IOPS

ServeTheHome showed Kioxia’s GP1 SSD running just above 10 million IOPS at FMS 2026, using PCIe Gen6 and second-generation XL-FLASH for a fast server storage tier.

Lenovo Whitsett AI Server Line Doubles Plant To 890,000 Square Feet
Chips & Semiconductors

Lenovo Whitsett AI Server Line Doubles Plant To 890,000 Square Feet

A Tom's Hardware tour details Lenovo's expanded Whitsett AI server line, including 890,000 square feet, up to 20 MW of power and liquid-cooling tests for GB300 and Vera Rubin racks.

Supermicro 160-Bay NVMe Server Reaches 20PB In 4U
Chips & Semiconductors

Supermicro 160-Bay NVMe Server Reaches 20PB In 4U

ServeTheHome examined Supermicro’s ASG-4116S-NU160R at FMS 2026, a single-socket AMD EPYC storage server with 160 U.2 NVMe bays, roughly 19PB to 20PB of capacity in 4U, and redundant 2.6kW power supplies.

Samsung $15 Billion Power Prepayment Proposal Ties Fabs to Grid Funding
Chips & Semiconductors

Samsung $15 Billion Power Prepayment Proposal Ties Fabs to Grid Funding

Reuters reports that KEPCO asked Samsung to consider a roughly $15 billion electricity prepayment as South Korea weighs grid funding for chip plants and AI data centres.

NeoMME Cuts Visual Retrieval Index Storage While Keeping Benchmark Accuracy
AI

NeoMME Cuts Visual Retrieval Index Storage While Keeping Benchmark Accuracy

NeoMME, released on Hugging Face, pairs 260M and 800M multilingual multimodal encoders with a retrieval design that cuts late-interaction index storage from about 1.5 MB to 6 kB per page.

AMD Prices 256-Core EPYC 9996 At $14,904 For Server Buyers
Chips & Semiconductors

AMD Prices 256-Core EPYC 9996 At $14,904 For Server Buyers

TechRadar reports that AMD’s 6th Gen EPYC 9006 “Venice” lineup includes a 256-core EPYC 9996 with 512 threads, 1GB of L3 cache, a 600W default power rating and a $14,904 list price for 1,000-unit orders.

AI Compute Shortage Shifts Chip Strategy Towards Efficiency
Capital & Policy

AI Compute Shortage Shifts Chip Strategy Towards Efficiency

Semiconductor Engineering analysed how AI data-centre growth is being shaped by foundry, memory, power and optical-link bottlenecks as chip buyers look for more tokens from scarce capacity.

Keep Reading

More Stories

Latest
Kepler Targets 2027 Production for HBM Replacement MemoryCloud & Data CentersOct 6, 2026Kepler Targets 2027 Production for HBM Replacement MemoryEE Times reports that Kepler Computing is preparing 3D ferroelectric memory for 2027 production, promising higher capacity and bandwidth per watt while limiting reliance on advanced-node lithography.Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceCapital & PolicyOct 6, 2026Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceYokogawa Engineering Asia has launched a Singapore center focused on OT cyber resilience, training, response planning and recovery coordination for Southeast Asia, Oceania and Taiwan.ClickFix Attack Uses Browser Cache To Hide Malware PayloadCybersecurityOct 6, 2026ClickFix Attack Uses Browser Cache To Hide Malware PayloadMicrosoft Threat Intelligence traced a ClickFix cache-smuggling method that preloads malware into browser caches, then uses file size checks and a pasted Run command to launch later credential-theft stages.VOA Tests Six-Month Startup Buildout Before Funding DecisionsFintech & Digital PaymentsOct 6, 2026VOA Tests Six-Month Startup Buildout Before Funding DecisionsTechCabal’s interview with VOA Venture Partners founder Victoria Olayide Adesanya describes a six-month build programme that lets the firm work inside African financial-infrastructure startups before deciding whether to invest.Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCrypto/Web3Oct 6, 2026Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCoinDesk reported that bitcoin stayed near $86,000 while the U.S. Dollar Index reached about 102.5, with U.S. rate expectations and European political risks strengthening the dollar backdrop.Google Freezes OSS Bug Bounty Reports After AI Submission FloodCybersecurityOct 6, 2026Google Freezes OSS Bug Bounty Reports After AI Submission FloodGoogle has stopped accepting new product vulnerability reports in its OSS VRP after invalid automated submissions swamped reviewers, while older reports and some Cloud VRP routes remain open.Fleuret AI Raises €4M For Continuous AI Pentesting PlatformCybersecurityOct 6, 2026Fleuret AI Raises €4M For Continuous AI Pentesting PlatformTech.eu reported that French startup Fleuret AI raised €4 million in pre-seed funding to develop an agentic-AI platform that turns penetration testing into a continuous security process.GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%Fintech & Digital PaymentsOct 6, 2026GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%A GFT Technologies analysis says AI-linked software documentation can cut maintenance effort and speed developer onboarding when knowledge assets stay synchronized with code changes.Schneider Electric Lines Up $22.6 Billion PTC DealAIOct 5, 2026Schneider Electric Lines Up $22.6 Billion PTC DealSchneider Electric plans to buy PTC in a cash transaction valuing the US engineering software provider’s equity at about $22.6 billion, adding product-lifecycle software to its industrial AI push.Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueCapital & PolicyOct 5, 2026Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueOla Electric founder Bhavish Aggarwal pledged 20 Cr shares to finance his participation in a rights issue that forms part of a larger ₹1,500 Cr fundraising plan.Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseAIOct 5, 2026Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseNatrona County trustees questioned whether teacher AI tools expose student data, even as existing district rules already ban unauthorized generative AI use by students.New Relic Reports US$18 Million GreenOps Savings After AI CertificationCloud & Data CentersOct 5, 2026New Relic Reports US$18 Million GreenOps Savings After AI CertificationA New Relic company news item carried by iTWire says the observability vendor has earned ISO/IEC 42001 certification, joined the EU AI Pact and reported US$18 million in GreenOps savings from more than 80 engineering initiatives.