r/DeepSeek • • 24d ago

News DeepSeek-V4.1-Flash Release (official)

414 Upvotes

It’s officially out and the prices have been updated.

///

Today, we officially release the DeepSeek-V4.1-Flash model. It is the smallest model in our new architecture family, with native multimodal visual understanding. The new architecture is designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models.

GPQA Diamond: 90.9
HLE: 36.8 (39.1*)
Codeforces (Rating): 3471
MathArena Apex: 65.6
Terminal-Bench 2.1: 90.6
Terminal-Bench 3.0: 30.0
Terminal-Bench 4.0: 31.2
DeepSWE v1.1: 74.2
ProgramBench: 20.3
NL2Repo-Bench: 65.4
CyberGym: 88.1
SEC-Bench Pro: 62.8
ExploitGym: 15.3
HLE (w/tools): 63.9
Automation-Bench: 54.8
Agents' Last Exam: 31.8
Chartography (w/tools): 78.9
BabyVision (w/tools): 89.6
ZeroBench-main (w/tools): 49.0
* Tested only on the pure-text subset of the HLE benchmark set.

API changes
DeepSeek V4.1 Flash is now available on the DeepSeek API with native multimodal support. Change the model name to deepseek-flash to call the latest V4.1 Flash model. The previous-generation models V4 Flash and V4 Flash Vision Exp have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash.

Meanwhile, extensive testing shows that V4.1 Flash now outperforms DeepSeek V4 Pro across performance, cost, speed, and total time, so we plan to retire V4 Pro in an orderly manner. After 12:00 Beijing Time on September 14, 2026, and until the future release of V4.1 Pro, all requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at the V4.1 Flash price.

API apricing adjustment
With the release of DeepSeek-V4.1-Flash, API prices have been reduced accordingly. For details, please refer to Models & Pricing.

///

Source:

https://api-docs.deepseek.com/updates/#deepseek-v41-flash-release


r/DeepSeek • • 12h ago

News DeepSeek presents an alternative to CUDA together with Huawei

Post image
572 Upvotes

DeepSeek and Huawei have joined forces to publicly present their alternative to Nvidia's CUDA ecosystem, releasing the code for TileLang, a high-level programming language adapted to Huawei's Ascend 950 chips, along with other software tools for core AI model tasks (computation and inter-chip communication libraries). The goal is for developers to be able to work with Ascend accelerators without having to rebuild from scratch the software infrastructure needed to train and run models, thereby making it easier to move away from CUDA in Chinese AI projects. In addition, both companies have developed a supernode of 128 Ascend 950 accelerators, and DeepSeek plans to install at least 160,000 of these chips in a huge data center in Inner Mongolia, in an effort to reduce dependence on Nvidia's hardware and software.

TileLang https://github.com/tile-ai/tilelang

  • 2026-09-30 — Ascend 950 backend: TileLang now officially supports Huawei Ascend 950 NPUs with native code generation, automatic scheduling and synchronization, SIMD/SIMT vector programming, etc. Explore the Ascend examples for GEMM, FlashAttention, and more.

What do you all think of this new ecosystem that DeepSeek and Huawei want to promote?


r/DeepSeek • • 9h ago

Question&Help What's the reason behind bans recently?

40 Upvotes

I have been using the app for role-playing over two years by now and I don't want to lose my chats :(

I did a lil scroll here and some people said that role-playing is no longer allowed in deepseek due to chinese restrictions. Is that true?


r/DeepSeek • • 6h ago

Funny The cache 💀💀💀

Post image
18 Upvotes

r/DeepSeek • • 8h ago

Resources Head to head: DeepSeek-V4.1-Flash vs Mistral-Large-3 — RuntimeWire

Thumbnail
runtimewire.com
14 Upvotes

r/DeepSeek • • 2h ago

Question&Help Deepseek V4 Flash Short-Comings... Or mine?

3 Upvotes

First I should clarify that I'm fairly new to using Cline. My prior experience is with Codex and before that VSCode with Co-Pilot, but mostly in a corporate context where I'm using expensive models and don't have to care as much about token costs.

I've been using Cline with VSCode for about a month so I can work on personal coding projects at home like I do at work. I use OpenRouter, and I've used various versions of Gemini 3 Flash and Deepseek for coding tasks. After writing .clineignore and some .clinerules I think I've shaved the waste when working with my projects and am getting decent cache hit rates.

There's no argument that Deepseek is far cheaper than Gemini 3 Flash, as whenever I ask Gemini to make a change it is usually quite fast but can spend a dollar on average for just 15 minutes of work. Sometimes this is an acceptable cost to me, as I've found it reliable and if I just want small things done quick it's a great solution. When I switch to Deepseek, I will spend only tens of cents on similar small projects.

The thing is: I think Deepseek could actually do those jobs for even cheaper, but it has a few problems that mess with its performance.

  1. Compared to Gemini, it screws up small changes far more often and corrupts files. I have added in clinerules guidance to use methods that are less error prone, but it has at times ignored this. Here I'm talking inaccurately injecting/replacing code to the wrong part of a file, creating duplicates and over-writing code that should not have been over-written. When it notices and tries to correct this, it often makes the situation worse. People seem to recommend Deepseek V4 Flash for the "Act" mode of cline, but that's my main reason for shying away from it.
  2. Seems to have problems with using Windows powershell that Gemini doesn't. I installed Powershell 7 to get around issues Gemini had with it, but I see a lot of Deepseek complaints of "truncated" outputs. Perhaps this is a *me* problem, not setting it up right... Maybe Deepseek always thinks it's working in Linux, because it keeps using commands not supported by Powershell? But Cline is set to use Powershell, and even when it realizes the command doesn't work on Powershell it will try to use the same failing command again later.
  3. In Plan mode, when I ask it to analyze a problem, I will be watching its thinking text and it visibly goes down fruitless rabbit holes, and even sometimes finds the correct root cause but then takes forever using failing methods to verify assumptions (probably because of #2 above). For instance, last night it correctly identified that the wrong version of a package was installed and that the code was using an older version of the API. I can see it has even identified it as a definite root cause, that it has read the publicized API to see that it mismatches, but wanted to confirm by reading the code of that package and kept hitting issues accessing the script itself. I had to cancel it after 2 minutes of it trying and failing and told it to trust what was publicized; only then did it give me a pretty good analysis.

Sometimes paying $1-2 instead of $0.10 to avoid wasted time and code corruption is more desirable, but I'd really like to use the cheaper model if possible. So my question is: Are these common experiences with Deepseek, and can they be mitigated?


r/DeepSeek • • 5h ago

Discussion Just chilling on deepseek random thought

3 Upvotes

Hey r/DeepSeek. I’m just t ahh you know a guy, and I’m a constant daily user. I recently got access to the voice gray-test, and after dealing with the V4.1 Flash mess, I just wanted to drop my honest predictions and thoughts on where this is heading. I posted a rough guess a few days ago while just chilling, but here’s the full breakdown.

  1. The "Black Lash" is real, and Flash models are the problem. Let’s talk about the elephant in the room. Flash models (like V4.1 Flash) are fast but dumb. They completely suck for storytelling and roleplay. Why? Because they lack deep reasoning. They use shallow, blunt safety heuristics that can’t tell the difference between a sci-fi plot and a real-world threat.

We’ve all seen the screenshots. The app tells you "Good morning. Let's chat," and then immediately suspends you for writing fiction. Getting banned for no reason just because your character is a villain is incredibly irritating. It kills creative flow and alienates the most dedicated users.

  1. The Core Demand: Fiction = Fiction, Real = Real The system revamp needs to be context-aware. Right now, the safety filters treat legitimate creative writing like a real crime. We need a balanced, two-tier system:

· Real World: Strict, zero-tolerance for actual harm. · Creative Fiction: Context-aware mode for verified adults that understands narrative framing, mature themes, and complex characters.

We need the AI to understand that fiction = fiction, and real = real.

  1. My Next Update Prediction: V4.1 Pro (Not V5) I don’t think V5 is next. My guess is we’re getting V4.1 Pro. And I’m predicting it will bring two major things:

· Full Release of Voice: Moving the current gray-test (Shell, White Wave, Starfish, Dark Tide) into a full, official rollout. · The System Revamp: A total overhaul to fix the context-awareness and stop the false-positive bans.

  1. The Ocean Naming Pattern and the Future Look at the industry pattern: Gemini 4 Argon, GPT-6 Astra, Claude Opus 5.5. DeepSeek is already using marine themes internally (Sealion, White Wave, Dark Tide). If they follow this branding, the next flagship has to be DeepSeek 5 Ocean (or Sea/Wave).

And looking further ahead to Summer 2027? I predict AI will have full YouTube and link access. Not just reading transcripts, but actually "watching" videos and using agentic "hands" to click, scroll, and navigate dynamic websites. The DSec infrastructure being built right now is the foundation for that.

I’m just a constant user chilling on Reddit, but I see the roadmap. The "Black Lash" from the community is going to force DeepSeek to fix the context-aware safety, and V4.1 Pro will be the proof.


r/DeepSeek • • 1d ago

Other I'm Part of the 1 Billion token club

Post image
104 Upvotes

Left all the US Frontier lab last week. 1B tokens would have costed me 200 to 1000 USD a month ago. Instead, I'm only paying $20.

Screw you Sam Altman and Amodei. Thank you China.


r/DeepSeek • • 15h ago

Discussion What does DeepSeek still do better than every other model for you?

12 Upvotes

Forget benchmarks. What's the one real-world task where you tried GPT, Claude, Gemini, Qwen etc. and still came back to DeepSeek?

Bonus points if you have a prompt that makes the difference obvious.


r/DeepSeek • • 17h ago

Discussion Deepseek Internal Monologue

Thumbnail
gallery
15 Upvotes

r/DeepSeek • • 19h ago

Discussion Agentic behavior

18 Upvotes

Since around a year ago, my business built our own agent software running on our FreeBSD servers. We selected to use the Anthropic models through their API and kept adding new models as they came out. The agent could route different types of task to different models. For basic tasks like reading, minimal text editing, and similar, it mostly selected Haiku. For planning, it used Opus, and later Fable. For larger workloads, it used Sonnet.

Roughly:

· Easy tasks: Haiku (with a 25% boost to Sonnet)

· Ordinary tasks: Sonnet

· Complex tasks: Opus / Fable

It was fully functional. We ran it like this all the time, only adding support for new models and occasionally extending our tools. But the API cost was kind of skyrocketing, and I have to say, Anthropic models are pretty bad for agentic work, especially low-level C and assembly.

Since a few weeks ago, we tried to replace the Anthropic API with DeepSeek, not for everything, it was more like a rest to get a better idea how the DeepSeek models are handling our type of work.

We have moved more and more to run via the DeepSeek models and right now, it almost feels like we’re running and using the models for free. It costs almost nothing. DeepSeek’s agentic behavior is much better then I expected, and the result from the DeepSeek Flash v4/v4.1 models is on pair or slightly better than identical work done with the Anthropic Sonnet 5/5.5 model.

Another issue with Anthropic, was the fact that they quite often flagged our work as a breach of some policy they had set up. While DeepSeek just gets it done without any policies refusal work...

Another very big issue with the Anthropic models are the fact that they are very, I should say extremely verbose. Their models output huge amounts of text, relative to the actual task.

We’ve run tests where the system prompt, tools, and complete agent software were identical between DeepSeek and Anthropic. In every single case, DeepSeek gave a much better result. I know that other people does not directly share our view regarding DeepSeek, or how worthless the Anthropic models actually are, compare to their cost. You are paying a premium that does not exist.

It almost feels like we were tricked into using Anthropic because it’s a Western company, while DeepSeek is from China.

TL;DR: DeepSeek has much better agentic behavior than the Anthropic models, and its cost is basically free compared to Anthropic.


r/DeepSeek • • 1d ago

Discussion Asked V4.1 Flash to extract all weapon assets from a Unity game. It didn't refuse!

Thumbnail
gallery
138 Upvotes

about 10 cents and 10 million tokens. Didn't refuse at all, went ahead without a single question.
Not planning to do anything with the models since i respect the devs, was just curious how deepseek would do.
Game: Gunman Contracts Stand Alone (VR Unity game)
37 models extracted with all textures.


r/DeepSeek • • 23h ago

Question&Help v4.1 flash + dsh web : are these good numbers ? Got lot of tasks done in 2 days.

Post image
8 Upvotes

r/DeepSeek • • 23h ago

Question&Help Anyone tried Deepseek 1-shot prompt?

4 Upvotes

As per title, has anyone tried one shot prompt to make something like how some YouTubers tested using Opus and Astra? Is this achievable via loop engineering and is the output comparable?


r/DeepSeek • • 1d ago

Question&Help Any tips to make DS 4.1 Flash stop yapping

4 Upvotes

The model just yaps too much rather than doing actual work. Like writes 5 types of test for a 10 line change where tests already exist, "thinks" for minutes and vomits nonsense


r/DeepSeek • • 23h ago

Question&Help What are the best providers for DS V4.1 flash? (on Openrouter)

2 Upvotes

Hey guys,

I recently set up Hermes agent on my VPS, for the model I chose Deepseek V4.1 flash since it's the sweet spot of cheap and smart, but I'm struggling with choosing a good provider and sticking with it. I want to set couple of providers on the whitelist and forget about it.

I wanted to use Deepseek itself, but the red warning beside it saying "this provider may use prompts for training and may retain prompt data" made me hesitate.
I used Deepinfra as well, it wasn't bad, but i saw some complains about it here.
All I know for sure is to put Open Inference on the blacklist and don't even look at it.

If you use this model, I'd be happy to know which providers worked best for you.

Thanks.


r/DeepSeek • • 23h ago

Question&Help ZDR via the official API?

2 Upvotes

Hi, does connecting an agent chat (via opencode, or some other open source harness) to an official DeepSeek API key allow ZDR?

If not, what other providers allows this?

I'm looking for something with a usage based billing, not a subscription.

Haven't used DeepSeek in the past but did a little research and it seems like the best value for money for my needs, so only the ZDR question is the only concern I have left.

Thanks!


r/DeepSeek • • 1d ago

Discussion No regeneration/edit limits on web

Post image
13 Upvotes

At least for now. I tested independently, > 20 regens/edits within 10 minutes in a single chat, no 'too many regen/edits' warning yet!


r/DeepSeek • • 1d ago

Question&Help Unable to use Deepseek in Android Studio due to JSON unknown roles

4 Upvotes

This is the error I am getting, is there a fix to this problem?

Model query failed: 400: Failed to deserialize the JSON body into the target type: messages[0].role: unknown variant developer, expected one of system, user, assistant, tool, latest_reminder at line 1 column 40207


r/DeepSeek • • 1d ago

Discussion Temporary 3 day ban was lifted, then got another 7 day ban

41 Upvotes

I posted last Friday that I was strangely banned from Deepseek for 3 days, and the only tie i could figure was that i was speaking russian with deepseek (i use mostly for code generation + documentation + other boring things). After reading the responses there, I now think perhaps the russian was just coincidental. the ban lifted 5 days ago, but was busy so did not use deepseek again until tonight. I used it in pretty boring ways (had it generate some static html, had it write a business email that i need to send monday, then tried to pressure test a recent decision). when i came back an hour later it said its banned for one week due to violations.

I did use a bit of russian during the chat, and the static HTML was lang ru, and all strings were russian. But other than that was very careful about using russian. I do not really understand the basis of these bans. The compelling idea i saw in the first thread was that high usage might be triggering it. Though i did not use it very much tonight.

At this point sadly i guess the platform is unusable for me. If i can't use it to have a brief (very boring) conversation or make a single static html page, i'm not sure how to use the platform.

I tried out chat gpt in the middle of this, and i was surprised how much i disliked it, and how different the reasoning was. this has been a surprisingly useful tool, any ideas?


r/DeepSeek • • 23h ago

Question&Help Can anyone sponsor me an AI coding subscription?

Thumbnail
0 Upvotes

r/DeepSeek • • 1d ago

Other Kosmora: a calm, good-looking AI coding agent.

2 Upvotes

Hey all. Meet Kosmora, an AI coding agent for Windows with its own desktop app.

The first thing I hope you notice is how it looks: dark, quiet, one accent color, nothing shouting at you. I tried to keep the surface simple and easy on the eyes, but there's a lot inside. You describe a task, the agent edits your code and runs your tests, and every change comes back as a diff you can review, accept or undo. Around that: a built-in terminal, source control with GitHub, parallel tasks, a different model for each job, MCP servers, skills, plugins, subagents, web search, workspace memory, diagrams and images right in the answers, live cost and context meters, and updates from inside the app. If you're new to all this, every settings page has a small "i" that explains the feature in plain words.

Bring your own API key (GLM, DeepSeek, or anything OpenAI/Anthropic compatible).

https://kosmora.dev

If something breaks, tell me what you did and I'll look into it.


r/DeepSeek • • 1d ago

News New Compiler based agent cuts costs by 2x and improves code intelligence (78.2% SWE-bench Verified @ 0.1¢)

Post image
1 Upvotes

r/DeepSeek • • 1d ago

Tutorial Use your code agent for graphic design

1 Upvotes

r/DeepSeek • • 11h ago

Discussion Cutting Ties with DeepSeek

0 Upvotes

DeepSeek had it coming. It's no longer a must-have for my work. I can officially cross DeepSeek off my list of AI tools for my projects.

What do you think about that?