News
CAPACITY TEST:

Hugging Face Adds 4-Bit Nunchaku Loading To Diffusers

Newsroom brief

Hugging Face added Nunchaku Lite support to Diffusers, letting developers load 4-bit diffusion checkpoints with `from_pretrained()` while using Hub-delivered CUDA kernels.

Verified against source materialEdited by SendTech Times Chips & Compute DeskSource: Hugging Face
Hugging Face Adds 4-Bit Nunchaku Loading To Diffusers
Image source: Hugging Face

Diffusers gets a 4-bit loading path

A July 23 Hugging Face blog post introduced a Diffusers integration for Nunchaku Lite, giving text-to-image developers a way to load Nunchaku-style 4-bit checkpoints through the standard `from_pretrained()` workflow.

The post describes modern BF16 text-to-image pipelines as often requiring 20-30 GB of VRAM, while earlier Diffusers quantization options such as bitsandbytes, GGUF, torchao and Quanto mainly reduce model weight storage.

According to the same technical note, SVDQuant runs the main diffusion-transformer layers with 4-bit weights and activations, or W4A4, so the denoising loop can use less memory and run faster.

Kernels arrive through the Hub

Nunchaku Lite loads without a custom pipeline class, separate inference engine or local CUDA compilation, according to the post.

The runtime patches relevant `nn.Linear` modules in a stock Diffusers model with SVDQ or AWQ linear layers before loading the checkpoint, and CUDA kernels are downloaded through the `kernels` package.

The integration uses `svdq_w4a4` layers for attention and MLP projections, with INT4 and NVFP4 variants.

A second `awq_w4a16` path covers precision-sensitive normalization and modulation components in model families including FLUX and Qwen-Image.

The hardware table says NVFP4 checkpoints require NVIDIA Blackwell hardware, including RTX 50 series, RTX PRO 6000 and B200 GPUs, while INT4 variants support Turing, Ampere and Ada devices such as RTX 30 and 40 series cards, A100 and L40S.

Benchmarks stay vendor-owned

According to Hugging Face's example, one Nunchaku NVFP4 transformer paired with a bitsandbytes NF4 text encoder generated a square 1024-pixel image in roughly 1.7 seconds on an RTX 5090, while peak memory was about 12 GB; the same paragraph placed the BF16 pipeline at about 24 GB.

In a separate RTX PRO 6000 Blackwell benchmark, the blog listed a BF16 baseline at 3.00 seconds for the full pipeline and 31.1 GB peak VRAM.

The blog's benchmark table listed Nunchaku Lite NVFP4 at 2.27 seconds and 20.6 GB peak VRAM, Nunchaku Lite with `torch.compile` at 1.68 seconds and 20.6 GB, and Nunchaku Lite with an NF4 text encoder at 2.29 seconds and 16.0 GB.

The post characterized the result as up to 50% lower peak VRAM, roughly 30% better latency and as much as a 1.8x speedup with `torch.compile`; those figures remain vendor benchmark claims rather than independent lab measurements.

Quantization workflow widens

The post also points developers to `diffuse-compressor`, a companion toolkit for calibrating, quantizing, packaging and publishing new Diffusers models.

The FLUX.2 Klein 4B example workflow says the inspection stage should identify 100 SVDQ targets, 3 AWQ targets and 6 dense outer linears before quantization proceeds.

The remaining boundary is hardware and model structure.

The compatibility note excludes Volta and Hopper GPUs from current 4-bit kernel support, and the generic Nunchaku Lite path does not infer architecture-specific fused rewrites such as combined QKV projections.

That makes the integration useful as a standard Diffusers loading route, while specialized engines may still matter when a model family needs fused modules beyond the generic layer replacement path.

The post did not disclose production customer names for the new Diffusers path, so deployment proof is still limited to the published benchmark examples.

Share this article
inXf

Related articles

More
Lenovo Whitsett AI Server Line Doubles Plant To 890,000 Square Feet
Chips & Semiconductors

Lenovo Whitsett AI Server Line Doubles Plant To 890,000 Square Feet

A Tom's Hardware tour details Lenovo's expanded Whitsett AI server line, including 890,000 square feet, up to 20 MW of power and liquid-cooling tests for GB300 and Vera Rubin racks.

Nvidia Japan AI Factory Plan Lists 140MW Capacity
Cloud & Data Centers

Nvidia Japan AI Factory Plan Lists 140MW Capacity

Tech Wire Asia reported that Nvidia’s Japan expansion includes a 140-megawatt AI factory using 13,750 Vera CPUs and 27,500 Rubin GPUs for the METI-backed FRONTia Project. The report did not name the site, capital cost, power supplier, construction timetable or first enterprise workloads.

Quantum Data Centre Push Shifts From Qubit Counts To Hybrid Workloads
Cloud & Data Centers

Quantum Data Centre Push Shifts From Qubit Counts To Hybrid Workloads

Data Center Knowledge reported that quantum computing work is shifting toward hybrid systems that connect QPUs with GPU and CPU infrastructure. Hyperion estimated the market at $1.4 billion in 2025 and projected about $3 billion by 2028, while the article did not identify production customers, signed deployment contracts or facility-level power loads.

Armenia AI Plan Centres On Blackwell Compute, Not Chip Fabrication
Cloud & Data Centers

Armenia AI Plan Centres On Blackwell Compute, Not Chip Fabrication

AI News places Armenia’s Firebird plan around imported NVIDIA Blackwell infrastructure, US export approval and power capacity rather than domestic chip manufacturing.

Anthropic Builds Custom Silicon Team For Claude Scaling
Chips & Semiconductors

Anthropic Builds Custom Silicon Team For Claude Scaling

The Next Web reported that Anthropic publicly confirmed an in-house silicon team for Claude, while saying AWS, Google, Nvidia and AMD hardware remain part of its multi-chip scaling strategy.

Cadence Adds AuraStack AI Agent For PCB And Advanced Packaging Design
Chips & Semiconductors

Cadence Adds AuraStack AI Agent For PCB And Advanced Packaging Design

The Register reported that Cadence Design Systems introduced AuraStack, an agentic AI system for PCB and advanced packaging workflows. Cadence cited a 15x productivity claim and named Nvidia among customers, but The Register did not include pricing, availability dates, full customer names or independent benchmark results.

SoftBank Earnings Shift AI Focus From OpenAI To Intel And Arm Costs
Chips & Semiconductors

SoftBank Earnings Shift AI Focus From OpenAI To Intel And Arm Costs

CNBC reported that SoftBank beat June-quarter profit expectations after a 1.3 trillion yen Intel gain, while OpenAI produced no investment gain and the AI computing segment widened its loss.

Nintendo Keeps Forecast After Earnings Beat Masks Switch 2 Unit Drop
Chips & Semiconductors

Nintendo Keeps Forecast After Earnings Beat Masks Switch 2 Unit Drop

CNBC reported that Nintendo beat fiscal first-quarter revenue and profit estimates while Switch 2 hardware sales fell 34.4%, leaving the company to rely on software, pricing and its unchanged annual outlook.

Keep Reading

More Stories

Latest
Hugging Face Hack Pushes AI Agents Into Cybersecurity SpotlightAIAug 8, 2026Hugging Face Hack Pushes AI Agents Into Cybersecurity SpotlightCNBC reported that Black Hat cybersecurity leaders treated the Hugging Face AI-agent breach as a turning point for governing autonomous cyber models rather than a one-off failure.Alibaba Tests Revenue Sharing For Commercial Qwen AI UseAIAug 8, 2026Alibaba Tests Revenue Sharing For Commercial Qwen AI UseAI News reported that Alibaba plans revenue-sharing terms for some commercial users of its next Qwen open-weight AI model, following a licensing pattern already used by Moonshot for Kimi K3.Meta Ordered To Fund $567M New Mexico Youth Mental Health PlanCapital & PolicyAug 8, 2026Meta Ordered To Fund $567M New Mexico Youth Mental Health PlanArs Technica reported that a New Mexico judge ordered Meta to provide $567 million for treatment, screening, awareness and prevention after finding that its platforms contributed to a public nuisance.Harvey Funding Talks Could Lift Legal AI Startup To $15.5B ValuationAIAug 8, 2026Harvey Funding Talks Could Lift Legal AI Startup To $15.5B ValuationSiliconANGLE reported that Harvey AI is seeking at least $500 million in new funding that could value the legal AI startup at $15.5 billion after annualized revenue passed $350 million.Vietnam Shows Shopee-TikTok Shop Race Tightening In Southeast AsiaScience & TechAug 7, 2026Vietnam Shows Shopee-TikTok Shop Race Tightening In Southeast AsiaTech Collective SEA wrote that Shopee’s Vietnam share fell from 61% to 53% between May 2025 and April 2026 as TikTok Shop rose from 33% to 44%, showing how social commerce is reshaping regional ecommerce infrastructure.China Opens Security Review Of Palo Alto Networks ProductsCybersecurityAug 7, 2026China Opens Security Review Of Palo Alto Networks ProductsChina's cyberspace regulator opened a security review of Palo Alto Networks products, with no named product line, technical flaw or decision timetable disclosed.AI Pioneers Split Over Risk As Compute Buildout AcceleratesAIAug 7, 2026AI Pioneers Split Over Risk As Compute Buildout AcceleratesData Center Knowledge reported that Geoffrey Hinton, Fei-Fei Li and Andrew Ng disagreed at Ai4 over AI risk, jobs, openness and regulation, leaving infrastructure investors to plan capacity amid unsettled deployment rules.SpaceX Asks FCC To Wind Down $4.5bn Rural Broadband SupportTelco & ConnectivityAug 7, 2026SpaceX Asks FCC To Wind Down $4.5bn Rural Broadband SupportLight Reading reported that SpaceX urged the FCC to sunset High-Cost rural broadband subsidies, while rural telecom and electric-cooperative groups said LEO satellite coverage cannot replace terrestrial network support.OpenAI Expands Free ChatGPT Access In GPT-5.6 RolloutAIAug 7, 2026OpenAI Expands Free ChatGPT Access In GPT-5.6 RolloutBleepingComputer reported that OpenAI is rolling out GPT-5.6 Sol for paid ChatGPT users and GPT-5.6 Luna for Free and Go users, pairing unlimited free text chats with a new reasoning control and additional safeguards for users believed to be under 18.JLL Data Centre Report Shows Middle East Pipeline Pause As FLAPD GrowsCapital & PolicyAug 7, 2026JLL Data Centre Report Shows Middle East Pipeline Pause As FLAPD GrowsData Center Dynamics reported that JLL's EMEA Mid-Year Data Centre Report 2026 put FLAPD live capacity at 3.8GW, while the Middle East had 2.6GW in development paused and 13.8GW in planning.AWS Adds Persistent Runtime Instances For Production AI AgentsCloud & Data CentersAug 7, 2026AWS Adds Persistent Runtime Instances For Production AI AgentsAWS announced runtime instances for Amazon Bedrock AgentCore Runtime, adding managed infrastructure for multi-agent workflows, shared sessions lasting up to 14 days and GPU-supported production agent deployments.AI Patch Study Keeps Humans In Vulnerability ReviewsCybersecurityAug 7, 2026AI Patch Study Keeps Humans In Vulnerability ReviewsThe Register reported that 1Password Off-by-1 Labs tested 6,080 AI-generated patches across six CVEs and found clean autonomous fixes in 26.0 percent of cases, leaving security teams with a supervision problem rather than a replacement for vulnerability review.