GitHub radar
Top 5 Hugging Face Models This Week
I will remember this week on Hugging Face for two things: a free model from Cloudflare sorts a ticket or an invoice into a number instead of a paragraph, and a video model from Lightricks holds one character across several joined scenes for the first time.
I installed LTX-2.5 through ComfyUI and tried to join several scenes into one clip in a row. In Lightricks' previous version, LTX-2, the model could only produce one continuous shot, but this version holds the character identity, the lighting and the overall style across several joined shots at once. The sound comes out together with the video in the same clip: the input is text, an image or a piece of an old video, and the output is a ready clip with its own audio track. I liked that the weights are open, and that I can fine-tune the model for my own tasks at home, not only reach it through someone else's service.
Why a vibe-coder should care
Before this, putting together a video with several scenes and its own sound meant stitching pieces from different programs. Here everything comes out of one model at once, and it can be installed and fine-tuned at home, instead of paying for a cloud service every time.
A Mac with 32 GB of RAM or a GPU with 24 GB of VRAM
Open on Hugging FaceI tried giving the model one invoice and two questions at the same time: whether it is paid or overdue, and whether the amount is above one thousand dollars. Clef-Flash from Cloudflare is a 9 billion parameter model that reads the record once and answers all the questions at once, instead of writing a paragraph that I would have to read through myself. On the test invoice for 1250 dollars with the overdue status, it gave ninety percent for the option overdue and only five percent for paid. I liked that it is compatible in format with the paid service Jev, and this version is open and free.
Why a vibe-coder should care
Tasks like sorting a ticket, an invoice or a comment by a fixed set of questions used to be solved only through a paid API. This model can be installed on a regular laptop, and the answer comes back as numbers instead of a text that I would have to read and sort out myself.
A laptop with 16 GB of RAM or more
Open on Hugging FaceI installed the quantized version through Ollama and gave the model the same kind of task I usually solve in Claude Code: find a bug from a ticket description in someone else's repository. Qwen3.8-27B from Alibaba is built on the Qwen3.5 architecture, and on paper the gain looks real: on Terminal Bench 2.1 it scores 73 points against 63.4 for the previous version, Qwen3.6-27B, and on computer use through a screen image it scores 84.3 against 63.9. The model sees text, images and video, and its context is 262 thousand tokens out of the box, extendable to one million. For Chinese models this is a familiar pattern: the benchmark numbers grow faster than the real quality in agentic work, and it is still far from the level I use every day in Claude Code.
Why a vibe-coder should care
For a routine task at home, or when you do not want to pay for a subscription, this model is good enough, especially if your own GPU is enough for it. For real agentic coding work I still stay on Opus in Claude Code.
A Mac with 32 GB of RAM or a GPU with 24 GB of VRAM
Open on Hugging FaceLast week I already tried this model and cut out a background in one request. In seven days the download count grew from 52804 to 90003: people keep installing this model. Qwen-Image-2.1 from Alibaba is a single model for drawing and editing pictures in the Qwen family, and in one request it accepts up to ten reference images, and an edit can be marked right on the picture with a circle or a mask, instead of a long text description. For the second week in a row this model holds a high place in the week's chart, and this does not look like a one-week accident.
Why a vibe-coder should care
For an article cover or a product card, a model with transparency and mask editing built in saves a separate step through a background removal tool. The downloads continuing to grow for the second week in a row show that this is not a one-week accident.
A regular laptop with 8 GB of RAM or more
Open on Hugging FaceI opened the ready demo space right in the browser and gave the model a floor plan with a question: which room is to the left of the kitchen, and it read the plan correctly. ZDTaichu5.0-9B from TaichuAI is a general 9 billion parameter model for understanding images, spatial tasks and acting as an agent, and by the vendor's own numbers it leads among models of about the same small size: 93.7 on the IFEval instruction test and 87.7 on the TAU2-Bench agentic test. The comparison here is only against other 9 billion parameter models, not against something at the level of Opus or Fable, which I use for real work every day. I liked that I could try it right in the browser, with no installation and no Ollama.
Why a vibe-coder should care
Reading a floor plan or asking a model where objects are on a picture - for a simple task like this at home it is enough, and you can try it for free right in the browser. For real agentic work I still compare the result with what I get from Opus in Claude Code, and the gap is noticeable.
A laptop with 16 GB of RAM or more
Open on Hugging Face▌ More finds