My day job involves a lot of contract work building all kinds of software around novel uses of LLMs. So we start with "Could LLMs do X?", and generally find this pattern:
- Holy shit, they can! This seems to work mostly!
- Oh shit, but not really. It makes inscrutable mistakes, and does so slowly.
- Actually, we should move 90% of this back into conventional code.
And we end up with the LLM basically being used to do intent/data extraction from text, and....well, that's it. LLMs are great for writing code, but at runtime it's difficult to find a use for them other than as a bridge from raw text to conventional software.
Which is actually fine, except that it's still expensive. On a lot of projects it just totally dominates compute costs just to do its (admittedly cool) trick.
So outside my dayjob, I have a travel planner project which helps people build custom Rick Steves-style european itineraries (tripsnek). As a fun experiment, I played around with a natural language interface a while back. It showed potential, but it also couldn't really be justified on a free app with no revenue. So that was that.
When the first decision models were introduced recently promising the LLM smarts at 1/10th the cost, it looked like the door opened back up, so I had to give it a try.
Edit: I just realized that, since these are so new, I should have included at least a brief explanation of what a decision model is. Essentially, they are LLMs that do not return any natural language. They only return structured decisions: yes/no, true/false, choose-from-a-list (e.g. pick from an enumeration). This is what allows them to be so efficient, and each decision is also associated with probabilities (which is something conventional LLMs are particularly unreliable at). I have intentionally avoided mentioning the company that launched the first one, since I think they are suspected of astroturfing marketing, and it doesn't matter anyway because there are already competitors open-weight alternatives since the cat is out of the bag on how they work (and that they work). Ok, back to the original post....
The app itself is essentially a conventional optimizer (a genetic algorithm), backed by lots of data about cities and travel times between them, and which can account for arbitrary user constraints and preferences. This last bit is what needs to captured by a user interface, and here I am taking a stab at a pure language interface as an alternative.
Here is an (ai-assisted) write up of the results, TLDR:
- Yes/no and choose-from-list parameters obviously worked incredibly well out of the box - preferred mode of travel, preferred pace, general interests, etc.
- Questions that required picking from very large sets - specific cities and sights, dates and ranges - were naturally much more challenging, requiring relatively complex regexes and text munging.
- It all netted out to a pipeline that matched, or even exceeded, a sophisticated conventional LLM provided with similar instructions, at roughly an order of magnitude lower cost...but also an order of magnitude greater code complexity (see figures).
Again, this is an initial stab, and I don't know if I'll keep pursuing it, but its already good enough that I am convinced it could be very good if I do and - just as importantly - its so cheap that its totally plausible for a free app sustained on a shoestring.
I'm not sure I even regard the additional code complexity as a burden. Sure, its more work, but it is also fully debuggable, optimizable, extendable code. If something goes wrong, the odds are now much higher I can actually fix it directly rather than with mystical incantations in a prompt.
Another dimension of this is that it offers the potential of capturing significantly more subtle aspects of user intent than are really feasible with a conventional UI. People have limits on how much time they are willing to spend figuring out how to twiddle knobs in a UI, which constrains how nuanced I can make my optimizer. As currently built, it cannot incorporate "the user kind of wants to go to Venice" because I don't know how I would build an elegant UI that would allow them to express it. Natural language could handle things like that organically, which is really neat.
Finally, I think I find decision models so refreshing because they don't have that dark cloud of generative AI. They do their job of answering questions about text humans wrote, then its back to software. It's just straight regular centaur shit.
Hope you all find this informative. You can try out the (experimental) interface here, and use the checkbox toggle to see a full drill down on all of the questions being asked and the answer that the decision model returns: https://tripsnek.com/describe/
AI Disclosure:
- This post = zero AI used
- Travel planner app = very little AI
- Jev pipeline = heavily AI assisted
- Results write up/figures = heavily AI assisted
Self-promotion disclosure:
- this post includes a link to my app