← All news·2026-06-30·2 min read

DeepSeek accelerates AI inference by up to 85% without requiring any special hardware

AI inference acceleration framework DSpark developed by DeepSeek: response times can be reduced by up to 60–85%, the number of requests per unit time has increased by up to 661% in testing. This solution is a response to the US restrictions on exporting high-tech chips.

aideepseekinferencechina

DeepSeek has developed the DSpark AI inference acceleration framework. The principle of operation is as follows: the draft model suggests several tokens at a time, and the main model checks their correctness simultaneously. Also, small batches of tokens are created and dynamic load balancing is performed depending on the reliability of the result.

Testing showed that the final response time from the user’s point of view has decreased by 60–85%, and the number of requests per unit time on the Gemma model from Google and the Qwen model from Alibaba has increased by up to 661%. You can try DSpark right now on GitHub and Hugging Face.

In order to reduce China’s computing power, the United States tightened the rules for exporting high-tech chips. With the help of DSpark, each percent of efficiency will decrease the dependency on American chips. Rather than closing the technical gap with new chips, it is possible to do this with clever programming, which is available to everyone without access to high-quality GPUs.

Source: the-decoder.com

Free course

Stop reading about AI — start building with it

The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.

Start free →
EAEvgenii Arsentev

Author

Evgenii Arsentev

PhD · Chief Executive Officer, digital health