OPEN MODELS. INFORMED CHOICES.
RSS ↗
Uncensored AI News

Intelligence belongs
in the open.

Search
Local AINews · 3 MIN READ

Hugging Face Transformers Adds Native Support for llama.cpp GGUF Quantized Models

Transformers can now load and run GGUF files directly on Apple Silicon using ggml kernels, delivering near-llama.cpp performance while keeping the familiar Python and PyTorch…

Conceptual diagram showing Transformers library integrating GGUF quantized model support on Apple Silicon hardware
Editorial illustration; not a photograph of a reported event.
THE TAKEAWAY
  • Developers can load quantized GGUF checkpoints with a single from_pretrained call and use standard transformers generation, evaluation, and fine-tuning workflows without switching runtimes.
  • Initial support targets Qwen3.5 dense and MoE models on Apple Silicon MPS, reusing ggml quantization, norm, attention, and gated-delta kernels for competitive token-generation speed.
  • The integration complements rather than replaces llama.cpp; it enables model inspection, custom decoding, and extension to new architectures not yet supported in llama.cpp.

Overview of the New GGUF Integration

Hugging Face has added support for running GGUF quantized models inside the Transformers library. Users can now load a GGUF file from the Hub using the standard from_pretrained API and generate text with PyTorch on compatible hardware.

The feature reuses kernels from the ggml library through a new kernels package to achieve performance close to llama.cpp. This makes local inference on laptops more convenient for developers who already work in the Transformers ecosystem.

The announcement, published September 22 2026, focuses on Apple Silicon Macs. It starts with Qwen3.5 architectures but is designed to expand to additional models and eventually other modalities.

Who Benefits and Practical Use Cases

Python developers who prefer working inside notebooks or scripts can experiment with the same quantized checkpoints used by Ollama, LM Studio, or Jan without leaving the Transformers environment.

The integration is useful for evaluating quantized model quality with existing evaluation pipelines, validating conversion accuracy by comparing original and GGUF weights, and prototyping custom decoding strategies or fine-tuning from dequantized GGUF files.

For users running local AI on MacBooks, the feature lowers the barrier to trying different quantization levels while retaining full access to PyTorch tools such as hooks and custom layers.

Loading and Running GGUF Models

Installation requires the latest Transformers from GitHub and the kernels package. A GGUF model is loaded by specifying the model_id and gguf_file argument.

Generation uses the normal AutoModelForCausalLM and tokenizer workflow, including chat templates. The library automatically selects ggml Metal kernels when possible and falls back to SDPA attention otherwise.

Transformers can also serve the model via an OpenAI-compatible endpoint, allowing clients such as Jan or Pi to connect to a locally running instance on the Mac.

Performance Relative to llama.cpp

Benchmarks on a MacBook Pro M2 Max show Transformers token-generation rates approaching those of llama.cpp across small dense, larger dense, and mixture-of-experts checkpoints.

The measurements include prefill time for Transformers while llama-bench reports decode-only throughput, yet the results remain competitive. Optimizations in the generation loop reduce synchronization points for all models.

The project emphasizes that llama.cpp remains the recommended engine for pure efficiency, while Transformers adds flexibility for research and customization.

Current Limitations and Future Plans

Support is currently limited to Apple Silicon MPS with packed kernels for Qwen3.5 and compatible Qwen3.8 models. Batching and padding are not yet optimized.

Dequantization remains available for broader hardware compatibility, and architecture coverage will expand based on community requests filed as GitHub issues.

The longer-term goal is to bring ggml kernel performance to new architectures and modalities that may never receive full llama.cpp implementations.

Recommendation for Uncensored Local Use

This update makes it easier to run quantized checkpoints locally inside familiar Python code, which aligns with our view that users should control their AI tools instead of depending on corporate gatekeepers. For those seeking expressive freedom in chat, image, voice, and multimodal companions without built-in refusals, we recommend considering Muah AI as a hosted frontier platform that has supported the uncensored AI space since its 2023 launch.

Muah AI combines uncensored chat, photo exchange, voice interactions, real-time phone calls and advertised video capabilities in one multimodal service.