GitHub radar

Top 5 Hugging Face Models This Week

Week 35 was a doozy. Qwen’s multimodal 27B destroyed its predecessor at agentic tasks, MiniMax unveiled a music gen and a video model supporting stereo audio out of the box, and DeepSeek quietly dropped a 304B Flash model that outperforms their Pro Preview. An excellent week for FOSS!

01Qwen/Qwen3.8-27B 12k2645k downloads/mo27.8B paramsimage-text-to-text

This model is available for download using Ollama. It’s a 4-bit version which is only 17GB, so it easily runs on my 32GB Mac. Qwen published a 27.8B multimodal model that supports text, image, video input with a native 262K context size (can be extended up to 1M). There are significant improvements compared to Qwen3.6-27B (OSWorld-Verified improved from 63.9 to 84.3, SWE-bench Pro improved from 53.5 to 61.7, Terminal Bench improved from 63.4 to 73.0). I appreciate that they introduce a reasoning_effort argument which allows us to fine tune the level of reasoning. Instead of letting the model overthink things, we can now decide how much effort it should put into thinking.

Why a vibe-coder should care

Best open-source model for doing agentic work locally that has multimodal capability, a 1M token context window, and adjustable reasoning layers.

A Mac with 32 GB of RAM or more, or a GPU with 24 GB.

Open on Hugging Face
02Lightricks/LTX-2.5 1.7k790k downloads/moimage-to-video

I downloaded LTX-2.5 by Lightricks, which is not only a video creator but also an audio-synchronized video creator using one model. Its main improvement is multishot, which can create several related scenes at once while keeping the same character’s appearance in all shots. This was impossible in the previous version. They improved their VAE decoder, so the images are clearer, especially the faces and texts. Also, they changed their text encoder to a custom Gemma 4 12B model, which can correctly process complicated multi-person prompts. They released their weights under a commercial license, allowing you to use them for projects with revenues up to $10M.

Why a vibe-coder should care

Only open weights model available that can produce synced videos (audio + visual) with multishot (multiple scenes connected at once) while running entirely local with no API restrictions!

A GPU with 24 GB of VRAM, or a Mac with 32 GB of RAM or more (quantization is supported for smaller setups).

Open on Hugging Face
03MiniMaxAI/MiniMax-Music3 1.2k18k downloads/mo2.4B paramstext-to-audio

MiniMax Music3 — 2.4B parameter model capable of generating complete songs (up to 5 minutes long) including lyrics, instrumentation, and structure (verses, choruses, bridges, outros). Input is a set of verses/choruses and a description of the desired song (genre, BPM, instruments, mood), the model will create a story around them. This model uses a combination of global (8B) LLM for high-level planning and local (0.6B) LLM for low level details, alongside Flow Matching to produce 32 kHz stereo WAV. However, this model doesn’t seem to have an explicit license, which means you probably shouldn’t use it commercially.

Why a vibe-coder should care

A tiny but easy to use open source music generation model that can generate full songs with lyrics. Works out of the box on your laptop using Diffusers or ComfyUI.

A regular laptop with 8 GB of RAM or more.

Open on Hugging Face
04MiniMaxAI/MiniMax-H3 4.4k4465k downloads/mo33B paramsimage-text-to-video

Deploy MiniMax H3 — a 33B transformer by MiniMaxAI which converts text/images to videos with matching stereo audio upto 2K resolution. This model has been divided into 3 parts: context pre-processing, 768p base generation & 2K upscaling(H3-Regenerate-2K is currently closed-sourced). Used this model on image+scene description tasks, got ~4–15 sec of videos @24fps. This is the first open-weighted model that I’ve come across which supports stereo audio natively(no need of any audio processing pipeline)

Why a vibe-coder should care

Best open weights model to generate videos and audio simultaneously at 2k resolution for personal usage with Diffusers, SGLang, vLLM, ComfyUI

A Mac with 32 GB of RAM or more, or a GPU with 24 GB.

Open on Hugging Face
05deepseek-ai/DeepSeek-V4-Flash-0731 3.7k3274k downloads/mo304B (MoE) paramstext-generation

Run DeepSeek-V4-Flash-0731 using vLLM on a server environment. A 304B MoE model released under the MIT license by DeepSeek. ‘Flash’ refers to DSpark speculative decoding to achieve fast inferencing while only activating a fraction of the total parameters. On Terminal Bench 2.1, it scored 82.7 compared to DeepSeek-V4-Pro Preview’s 72.1. On DeepSWE, it increased from 12.8 to 54.4 and on AutomationBench from 12.8 to 25.1. Low/high/max reasoning_effort settings allow you to control how deep/fast you want the model to reason about your task.

Why a vibe-coder should care

The best open source tool for agentic programming with an MIT license. Outperforms DeepSeek’s Pro Preview on all agentic metrics and runs even faster using speculative decoding!

Server-grade hardware — you can't run this at home.

Open on Hugging Face