GitHub radar

Top 5 Hugging Face Models This Week

This week on Hugging Face I remember for two things: an image model that finally gives a transparent background by itself, and the 27 billion parameter reasoning model I praised last week picked up another one and a half million downloads in seven days.

01Qwen/Qwen-Image-2.1 2.5k53k downloads/mo7.1B paramstext-to-image

I installed the GGUF build from unsloth and tried cutting a background out in one request. Qwen-Image-2.1 from Alibaba draws with transparency for the first time in this line, so I got a PNG with an alpha channel instead of a picture I had to clean up afterward. In one run it accepts up to ten reference images, and an edit can be marked right on the picture with a circle or a mask, instead of a long text description. The maximum resolution is 2752 by 1536 pixels, which is enough for an article cover or a product card.

Why a vibe-coder should care

Before this, getting an image without a background meant running the result through a separate background removal tool. Here the transparency is built into the generation itself, and the result is ready for a product card right away.

A regular laptop with 8 GB of RAM or more

Open on Hugging Face
02Edge0/Audio8-ASR-Infinite 1.2k19k downloads/mo4.1B paramsautomatic-speech-recognition

I ran it on a live microphone stream and left it working without stopping it. Audio8 ASR Infinite from the Edge0 team is a speech recognition model that listens continuously, and it has no limit on recording length because its memory and delay do not grow no matter how many hours it keeps running. It can tell a thinking pause apart from stuttering and from a real end of a sentence, so it does not cut a person off in the middle of a word. On the English LibriSpeech test it makes about one error per thirty three words, and on the Chinese AISHELL-1 test it is under two errors per hundred.

Why a vibe-coder should care

Before this, streaming recognition with constant memory and delay and no time limit was mostly available only through paid APIs. This one has an Apache 2.0 license, and it can run on your own machine, for example for live captions or dictation inside your own tool.

A regular laptop with 8 GB of RAM or more

Open on Hugging Face
03XingChen-AGI/Xing4.0-29B-A4B 1.8k45k downloads/mo29B paramstext-generation

I installed it through vLLM and gave it a task to read someone else's repository and find a bug from a ticket description. Xing4.0-29B-A4B from China Telecom is a model with 29 billion parameters, of which only 4 billion activate per token, so it is noticeably lighter to run than a typical model of this size. Its context is 256 thousand tokens natively and extends to 512 thousand, so a whole small project fits in at once. On the SWE-bench Verified benchmark, where a model fixes real bugs in open repositories, it scores 75 out of 100.

Why a vibe-coder should care

Before this, an agentic model with this context and this result on real bugs was mostly something you would look for only among closed services. Here the weights are open under Apache 2.0, so it can be installed as a local helper for an agent that works with code.

A Mac with 32 GB of RAM or a GPU with 24 GB of VRAM

Open on Hugging Face
04prism-ml/Ternary-Bonsai-2-27B-gguf 2.2k3344k downloads/mo27B paramstext-generation

A week after I first installed it, I checked the download counter and it had grown by almost one and a half million, past three million in the past month. Ternary-Bonsai-2 from the Prism ML team is the same 27 billion parameter reasoning model that Qwen built, only the weights are recalculated so the file takes 5.95 gigabytes instead of 54. Across fourteen reasoning benchmarks it holds 98.2 percent of the full version's result, and on math and coding the difference is within the margin of error. I ran it through llama.cpp on my own Mac and got about 47 tokens per second, which is faster than I can read the answer as it streams in.

Why a vibe-coder should care

A week ago, a reasoning model like this could not run without a Mac with 32 gigabytes of memory or a GPU with 24. Now a regular laptop is enough for it, and this is not a one-time spike: it was downloaded another one and a half million times over the past week.

A regular laptop with 8 GB of RAM or more

Open on Hugging Face
05XingChen-AGI/TeleOCR 62728k downloads/mo1.2B paramsimage-text-to-text

I photographed a receipt with my phone at an angle, crooked and slightly crumpled, and fed it to the model with no preprocessing at all. TeleOCR from XingChen-AGI is a document recognition model that weighs only 1.2 billion parameters, and it straightens out the skew and the creases on the page by itself, without a separate alignment step before recognition. It pulls not only text out of a photo but also tables and formulas, keeping their structure instead of just dumping plain text. On the OmniDocBench benchmark it scored 96.87 points and beat the specialized MinerU 2.5 Pro on this test.

Why a vibe-coder should care

Before this, recognizing crooked photographed documents with tables required a separate alignment step before OCR. Here the whole path from photo to text and table is one model, and it is light enough to install locally.

A regular laptop with 8 GB of RAM or more

Open on Hugging Face