r/technology • • 21d ago

Artificial Intelligence Bernie Sanders proposes 20 year prison sentence for AI devs who plow ahead with Artificial Superintelligence plans - penalty on par with illegally developing rogue nuclear weapons

https://www.tomshardware.com/tech-industry/artificial-intelligence/sanders-proposes-20-year-prison-sentence-for-ai-devs-who-plow-ahead-with-artificial-superintelligence-plans-penalty-on-par-with-illegally-developing-rogue-nuclear-weapons
48.0k Upvotes

1.8k comments sorted by

View all comments

Show parent comments

20

u/StrangeWill 21d ago edited 21d ago

As someone that has been a developer since 2003 (and runs an IT services company), ehhhh. That's not what diminishing returns means. Just because it crossed some threshold doesn't mean that it isn't taking an absurd amount more of resources to get less of a gap closed.

And ultimately: they're useful, but they still require a ton of guidance, every time the newest frontier model comes out it'll still randomly ignore instructions, it'll still implement far from optimal solutions, and when you correct it "you're absolutely right". I've had to step in on dumb decisions Fable has made on behalf of the team and resulted in significant improvements on performance, scalability and general performance of the solution -- on its own it'll still do dumb things.

I'll watch AI influencers wax poetic about whatever new method will "fix" that, then watch software released by those following those methods be even worse than before.

They're not knowledge databases, and they will not, by design, find best (or in some cases even good) solutions, in situations where we need solutions that good.

Sure, not every software stack needs solutions that good, but when you do, you see how little progress they're making with every model release.

and they still restrict Fable model usage to only 50% of the total usage you get from their subsidizes plans.

That's more of a financial decision then a resource limiting one, otherwise it's even more of a money pit.

The only thing that seems to be increasing in a way that is honestly impressive to me is local model performance, and again, that's because the frontier models are hitting diminishing returns the local models haven't hit yet, it allows them to close the gap.

3

u/michaelfrieze 21d ago edited 21d ago

That's fair regarding the definition of diminishing returns. I wasn't arguing that capability is scaling linearly with compute.

Where I disagree is the idea that this means the improvements aren't significant in practice. A relatively small improvement can be the difference between a model failing often enough that I can't trust it and succeeding often enough that I can actually incorporate it into my workflow.

That's what I've experienced over the past year. Fable and Astra still make mistakes, but far less often. With a good AGENTS.md/CLAUDE.md, project-specific skills, testing, and clear constraints, I don't find that they ignore my instructions very often. I'm increasingly reviewing mergeable code rather than rewriting what they produce. Also, I don't get the "you're absolutely right" thing anymore. That happened with older Opus models, but it's not something I have experienced recently.

I also think you're focusing too much on code generation itself. These agents can use terminals, browsers, computer use, GitHub, test runners, documentation, etc. They can investigate a codebase, reproduce a bug, implement a fix, run it, and iterate. I also use them constantly for smaller things: investigating issues and PRs, Git, commit messages, onboarding, one-off scripts, and internal tools. Those improvements add up.

Internal tooling and experimentation are especially important. There are plenty of ideas I previously wouldn't bother trying because they weren't worth days of engineering time. Now I can try several approaches quickly and throw away the ones that don't work.

As for AI influencers producing bad software, I don't know what examples you're referring to. You still have to be a skilled developer and know how to architect, constrain, test, and review what the agent produces. Bun's recent Zig-to-Rust work is a good example of experienced engineers using these tools for serious work: https://bun.com/blog/bun-in-rust

And I don't think "it won't reliably find the optimal solution" is a useful threshold. Human developers don't reliably find optimal solutions either. The question is whether it can produce good solutions consistently enough that reviewing its work is faster than doing everything myself. Increasingly, the answer is yes.

We're also seeing frontier systems tackle genuinely difficult and novel problems, including OpenAI's recent Navier–Stokes work. So I don't buy the idea that they're somehow incapable by design of producing novel solutions.

Diminishing returns at the resource level and rapidly increasing usefulness at the application level can both be true. I'm arguing that the latter is happening right now.