Long-time senior product designer here. I've worked with engineering teams of every size, including my own venture-backed startup (pre-LLM craze). Handoff always made sense, standups made sense, and I could ask questions whenever I needed to. You don't need to be super technical to be a successful product designer. Most of what I've ever had to ask is stuff like "hey, can we get that data?", "can we surface this stat here?", "can this menu do this transition?", "is this a native component or custom?", "what happens if this call fails?", "is this secure?", or "why can't this section load faster?" Very crude, I know. Maybe UX design actually teaches people more of this stuff, or maybe I'm literally insane for saying you don't need to be technical, and dev Reddit is about to eat my soul. I'm not very plugged into design circles, so it's possible I've just been doing this wrong my whole career and somehow bluffed my way through.
My last contract ended 6 months ago, so I don't have a big Slack group to throw questions at anymore. And asking AI isn't a great substitute. It's great at coding, but ask it about almost anything else and you get confident answers that fall apart when you push back (We all know this but not my point)
I hate the term vibe coding, but I've vibe coded tons of iOS apps and a couple web apps. Shit goes wrong all the time, but it's fun, and it's kind of a dream come true for a career product designer sitting on years of mostly crappy ideas.
Bit of a rant, but I wanted to give some background on where I'm coming from.
So my fucking question:
Why doesn't Claude (or ChatGPT) run a second pass on its responses before sending them? Specifically summarizing. When I get a wall of text and type "TLDR," the short version is almost always good. In my own apps I just add a step that sends the output back to the model with a new prompt, and it works really well.
Telling it to be short up front doesn't work for me. Asking for a TLDR after does. So why not automate the part that works?
And before anyone says "the model already reasons before it answers": maybe, but the replies still come out long. I'm talking about a pass that specifically makes them shorter.
My guesses:
- Cost. Adding a second model call to every reply at their scale still adds up, and these companies are already burning billions.
- Speed. You'd have to wait for the long answer to finish and then wait again for the short one.
- They expect you to just tell it to be concise. I literally have "concise" in my personal preferences and it still doesn't stick. There also used to be a style setting for shorter replies. It didn't really work, and now it's gone.
Sales: I don’t think the answer is that It serves their business interests so that enterprise teams will buy more if they use all their tokens is an applicable because I’m just talking about the Max subscription plan
If you worked at OpenAI or Anthropic and someone pitched this, would people laugh it off because of cost? Or is there a reason I'm not seeing that makes it harder than it looks?