MiniMax-H3 — MiniMaxAI
Multimodal · Audio-Video Generation
Synchronized Audio-Video
Likes: 2,418 · Trending Score: 2,318
The week's defining release. MiniMax-H3 generates natively synchronized audio and video from text, images, or existing video in a single diffusers pipeline — audio-video-generation in one forward pass, not a concatenation of separate models. ComfyUI integration followed within 48 hours, making the first truly open synchronized audio-video generation model immediately composable in local workflows.
View on HF →
DeepSeek-V4-Flash-0731 — DeepSeek
Text Generation · 1M-Context MoE
Efficient Long-Context LLMs
Likes: 2,452 · Downloads: 433K
The fast variant of DeepSeek-V4: 284B parameters with only 13B activated, supporting 1M-token contexts at 10% of DeepSeek-V3.2's KV cache through hybrid Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA). Dropped July 31 and immediately topped the download charts, with Unsloth's GGUF quantization arriving within hours of the official release.
View on HF →
Kimi-K3 — Moonshot AI
Multimodal · MoE Vision-Language
Efficient Long-Context LLMs
Likes: 10,093 · Downloads: 1.1M
Now in its second week of dominance, Kimi-K3 has crossed 10,000 likes — a milestone no other open model reached this quickly. Its 2.8T-parameter MoE architecture with 104B active parameters, native vision, and a 1M-token context window continues to drive massive community deployment as quantized builds proliferate across hardware tiers from data centers to laptops.
View on HF →
MiniMax-H3 — Comfy-Org
ComfyUI Workflow · Audio-Video
Synchronized Audio-Video
Likes: 724 · Trending Score: 685
The ComfyUI integration for MiniMax-H3, released within two days of the base model. Its rapid appearance demonstrates the community's capacity to make frontier multimodal generation immediately composable in local pipelines — the same pattern that makes every major open release simultaneously a lab artifact and a user tool.
View on HF →
DeepSeek-V4-Flash-0731-GGUF — Unsloth
Quantized GGUF · 1M-Context
Open Local Inference
Likes: 492 · Downloads: 111K
Unsloth's same-day GGUF quantization of DeepSeek-V4-Flash-0731, published within hours of the official release. The 111K downloads in days underscore the point: the gap between a frontier model release and its local-deployment form is now measured in hours, not weeks. Million-token context is moving to consumer hardware faster than any prior generation.
View on HF →
Unlimited-OCR — Baidu
Multimodal · Long-Context OCR
Open Local Inference
Likes: 3,899 · Downloads: 2.7M
Now with 2.7M downloads, Unlimited-OCR remains one of the most-deployed models on Hugging Face. Its Reference Sliding Window Attention eliminates the growing memory cost of long-document transcription, enabling thousands of pages in a single forward pass. A concrete illustration that efficiency — not just capability — is what drives real-world adoption at scale.
View on HF →
Inkling-Small — Thinking Machines
Multimodal · MoE VLM
Open Local Inference
Likes: 303 · Downloads: 15.5K
A compact multimodal MoE from Thinking Machines, supporting image-text-to-text and audio-text-to-text tasks under Apache 2.0. Inkling-Small represents the proliferation of capable small multimodal models — the frontier's capabilities, compressed into locally deployable footprints that democratize vision-language reasoning beyond the cloud.
View on HF →
LFM2.5-2.6B — Liquid AI
Text Generation · Edge Multilingual
Efficient Long-Context LLMs
Likes: 245 · Downloads: 47.4K
Liquid AI's LFM2.5-2.6B is a genuinely small model — 2.6 billion parameters — supporting 19 languages including Arabic, Chinese, Japanese, and Korean for edge deployment. As frontier capabilities flow downward through the size hierarchy, models like LFM2.5 mark the point at which multilingual intelligence becomes a commodity for embedded and offline use cases.
View on HF →