r/LocalLLaMA • u/No_Algae1753 • 16h ago
Question | Help Is there any way to improve creative writing for Local Models (Qwen)?
I wanted to know if theres anything that can improve creative writing for our Local Models? I specificly am asking for qwen models as they are way better when it comes to researching and writing html files compared to gemma / muse (which I know are better at creative writing). Im currently using qwen 3.8 flash next at q4 with llama.cpp
10
u/CatchDublinSurprise 14h ago
Why are we hung up on using only Qwen models (or, one specific Qwen model)? Use the model that's best for whatever it is that you're doing in the moment. Writing the HTML? Use the best coding model you can. Writing creative content? Use the best creative writing model you can. Coding agentically? You can probably figure out a way to have a creative sub-agent so you don't even need to think about it.
That being said, it's worth first experimenting with different creative models and prompts to find the combination that best fits your writing style. "Write in the style of a modern [insert your favorite author]" will go a long way. Keep it simple, though, and remember that LLMs don't understand negation as well as people seem to think they do. Things like "avoid purple prose the following phrases are banned [insert long list of slop phrases]" is more likely to make the problem worse, not better, since it draws the LLM's attention to those phrases.
Also, if you care about final quality, expect to do a manual final pass (or hit "regenerate" a bunch of times) until you're happy with it. AI is not yet capable of knowing whether its writing is good. Quality control by a skilled human writer is still necessary if quality is a concern.
1
u/Solembumm3 13h ago
Qwen flash next is a lot better at writing than Gemma 4 26-31b tunes.
2
u/CatchDublinSurprise 12h ago
In a very small comment you've packed a lot that merits pushback:
- Most finetunes out there are probably not worth your time. Anyone can make a finetune (and lots of people try), but very few people have the knowledge, datasets, and processing power to make a finetune that is superior to the original for a given task. My gratitude goes to the rare individuals who possess that trifecta.
- According to which human-scored benchmark does Qwen3.8-flash-next beat Gemma-4-31b in creative writing? Gemma-4-31b punches way above its weight in creative writing, and Qwen model training is highly focused on agentic and coding. Smaller models can't be good at everything.
- Having only 4b active parameters makes Gemma-4-26b-a4b a pretty lousy choice unless you're limited by hardware. Between the two, Gemma-4-31b is a much better choice for creative writing. I wouldn't cavalierly lump them together.
- I didn't actually say anything about Gemma. If you don't have the hardware to run GLM-5.3-flash or DS4-flash locally, there are other options (like the Hemingway finetune of Qwen3.6-27b mentioned elsewhere in the comments) that are worth trying (besides Gemma-4-31b).
2
u/Solembumm3 12h ago
Personal use. All fours main small gemma and qwen base model + tunes language capabilities (tendency to use very predictable small llm patterns) and context understanding are wastly lacking, in comparison to modern qwen next and dv4 flash models, that run in comparable speed category through mmap.
Hemmingway-1 is good for small tasks, like singular image gen prompts/descriptions for Krea, but on anything bigger have all the same problems.
18
u/Diecron 16h ago
I'd recommend looking at the router mode in llama.cpp to define the models you have, then using them directly per task. e.g. it would auto unload + load gemma for creative tasks, then back to qwen for others. I personally also have a router service running for short lived embeddings and reranking models (so they don't hold VRAM resident for too long).
8
2
u/No_Algae1753 16h ago
Yeah the problem is that it would take ages to recalculate the whole prompt. So switching is not an option sadly
2
u/Diecron 15h ago
Can you separate them? e.g. have the qwen write the prompt for the gemma model so you're not rerunning the entire prompt processing?
1
u/No_Algae1753 10h ago
How would that be possible? with qwen3.8 flash next i dont have space left for another model
1
u/Diecron 5h ago
from your other posts it sounds like it's in a tool capable harness, so the regular approach would be to have the main agent start a subagent with it's own prompt. Then, qwen can use whatever model's output verbatim instead of creating it itself. But, thinking about it you may not be able to have the model adequately summarize if it would otherwise need to see the whole context - in that case it should launch as a 'fork' of the conversation, but you're then right you'd hit a harsh PP penalty on that subagent.
2
u/GrungeWerX 15h ago
Gemma has a lot of AI slop in its writing. I know people like it, but it’s way worse than Qwen. I’ve heard of Hemingway, but never tried it.
Honestly, I’d recommend not using AI for the actual writing, but for the drafting. Creating treatments and such. You’ll get much better productivity out of them, they’re great for that sort of thing. They’re also good at critiquing your work.
You can use agents to help build entire pipelines where you can draft an entire plot in one go. You’d have to teach it story structure and your method, but you’d essentially have an assembly line and could workshop multiple scripts increasing your productivity to insane levels.
Remember: 90% of writing is rewriting. So don’t let AI take away your skillset. It’ll never be as good as you will with a lot of practice.
1
u/toothpastespiders 9h ago
Gemma has a lot of AI slop in its writing.
I often push the output through a de-slop phase. Gemma's surprisingly good at removing slop as well. Just not within it's own reasoning block. Or often with longer context. But I've had fairly good results having it write, then loop through it a couple paragraphs at a time while hooking into an additional tiny model trained to detect LLM tells and a prompt spelling out my style guide.
1
u/GrungeWerX 8h ago
Yeah, I had plans for that as well - basically a refiner - but I haven't had time to implement it yet. The original plan was testing both gemma 4 and Qwen 3.5 as refiners to see which handled the task better. 3.5 because it's actually really good for longform writing tasks and I have a very detailed treatment system I'm building a pipeline for. Also, 3.5 scored significantly higher than gemma 4 on the EQ bench - lower slop and just above one of the deepseeks.
Wish we could get something on par with Gemini 2.5 Pro 3-25 edition. No closed source model could match it on writing tasks I gave it and it's the only model that could successfully complete my complex treatment test.
0
u/silenceimpaired 15h ago
Yeah, AI is great for brainstorming and editing spelling and grammar… and that’s about it.
-1
3
u/msalsas 12h ago
Have you tried a Hermes fine-tune? I picked a Hermes 8B for a persona-driven project (it has to hold a character like a stoic philosopher or a poet through long conversations) precisely because the Hermes tunes are noticeably better at voice and creative prose than the base instruct models at this size.
3
u/Simpsator 10h ago
So maybe not quite what you're looking for but I'm starting to use a bit of a mixed model approach for script based creative writing (for a video game). My main harness is Claude, and it excels at keeping long story context, historical and character knowledge, but suffers when it comes to prose. Unfortunately my local options are a bit a bit limited by 12gb of vram, but I've found TheDrummer's fine tunes Cydonia 24b and Magidonia 24b are excellent at prose, but mediocre at long context. So recently more my hybrid approach is to use the larger model to prompt the small models (at a high temperature) for prose and "creative spark" for lack of a better term, basically generate a lot of iterations then let the larger model pick and choose and stitch it together for my approval. You might be for to do the same with QWFN and smaller prose-oriented models if you can fit both on your hardware. Take a look at TheDrummer's fine tunes in general. They are leagues ahead in creativity and prose than the base models.
5
u/XiRw 16h ago
I never really tested this so this could sound like complete utter shit but you can have one of your Qwen models do a web search on a book, pdf, skit, whatever (or upload one) that you really like and ask it to emulate a type of character or just the vibe of the book itself . That way it gives it clear instructions on which direction to go instead of just relying on the model itself. I might try this myself now to see if it helps
3
u/silenceimpaired 16h ago
There is a way… fine tuning: https://huggingface.co/Altworld/Hemmingway-1
2
4
u/Responsible_Pin_8965 16h ago
for rp with qwen ive found adding a few lines in the system prompt about describing thoughts and sensations in detail makes the creative output way more consistent.
3
2
u/a_beautiful_rhind 13h ago
Something else in that weight class? GLM 5.3-flash? Supposedly inkling was good at both but support for it is nonexistant.
2
u/toothpastespiders 9h ago
I've found RAG can be really helpful for this when carefully applied. But it's a lot of work to build up a good RAG system and populate it. I have a sort of brainstorming phase where a variety of methods are used to group together associations based on analyzed elements within the context. So frontend -> breaks down the components of the subject matter -> graph-hybrid RAG -> separate LLM call to streamline the jumble of semi-connected results -> back to the model to continue its work while leveraging the new associations.
The nature of the associations and methodology used there varies a lot on a task by task basis. Ranging from an attempt to project probable paths after the end of a chapter in a novel it's reading through to more of a surreal daydream type concept. I think the end result is partially logical and partially just a sampler concept pushed up in intensity.
The RAG system itself is basically categorized data from anything and everything that catches my attention. Real pain to put together back in the day. But at this point I'm betting you could vibe code something along that line pretty easily.
1
u/infieldmitt 15h ago
i find with QFN, the more thoroughly i prompt it, taking care to be clear about everything [character manifest is here, sample prose is here, etc], and giving a very rough outline of what i want to see, it'll do far better than when i just say "be creative".
2
1
u/cromagnone 14h ago
Pass the decision output from Qwen as JSON to one of the better models for writing, along with a prompt that every JSON item must be semantically included in the output.
1
u/Egor4more 14h ago
In my experience Qwen is good at not making stuff up and staying on topic but not good at playing a certain role, it’s a nerd. If I need a specific vibe from it, control vectors usually work better than prompting. Considering you are on llama.cpp, it should be pretty easy to try out representation steering. Though whether this will work for you or not depends on your task and I haven’t tested control vectors with 3.8 flash next. On 27B works very well though
1
u/Chida82 13h ago
This is interesting. How specific can you get with control vectors for writing style?
Are you steering broad traits like descriptive vs concise, or can you actually push things like dialogue style, pacing, atmosphere, etc.? I'd love to hear how you build/test the vectors for this.
1
u/Egor4more 12h ago
My experiments were more about single-character control rather than about overall book-writing styles but I believe vectors work for both. Some of the ones I tried are: aggression, agreeableness, cognitive clarity, neuroticism (emotional stability), trust, optimism, verbosity etc.
Atmosphere can be steered for sure, expressiveness vs dryness and similar should work as well, but be aware that if you steer things like "happiness", if your single model writes dialogues for several characters, that trait will be applied to all of them. If you have any specific vectors in mind, I could generate a couple for qwen3.8-27B (or some gemma4 model) and send the results.
I generate and evaluate the vectors on the llama.cpp fork I built specifically for that. It takes two persona descriptions as input, positive + negative (e.g. "you write vividly and express your thoughts through long beautiful sentences" + "your texts are always straight to the point without unnecessary complexity"). It then takes under a minute to generate the vector and (optionally) several minutes to evaluate that vector. I try to keep it up to date with upstream and it's compiled for most platforms so feel free to give it a try (https://github.com/Egor4More/ucvg.cpp)
1
u/jacek2023 llama.cpp 13h ago
I use Qwen on 4x3090 and Gemma/others on 2x3060 at the same time. So Qwen is telling Gemma what to do.
1
u/Chida82 13h ago
That's a neat setup. What do you actually pass from Qwen to Gemma?
Does Qwen produce a compact outline/instructions and let Gemma write from that, or do you pass some of the original context too? I'm curious how much context the creative model really needs.
2
u/jacek2023 llama.cpp 12h ago
pi coding agent runs shell commands and python scripts, I have whole set of tools to manage llama.cpp
1
u/ArtfulGenie69 9h ago
You don't want to use qwen for creative writing you want to use Gemma 4. The best I've found is Gemma 4 styletuned and if you need it huggingface has a version that is hereticed as well. Running that model with DRY enabled makes it a good writer.
1
u/marhalt 4h ago
I do a lot of work with creative writing with models. I second what other people said - split your pipeline, with whatever agenting / writing html work can be done by Qwen, but then use another model to do the creative writing. GLM 5.2 is by far the best writer, GLM flash can do, and Deepseek Flash may be what a good one-model-does-both compromise if you can run it. It's not GLM, but it writes reasonably well with a good prompt, especially if you enable thinking. Otherwise, it really depends on the kind of creative writing you do. There are some fine tunes (hermes, etc...) that can do some things, but generally nothing that will touch the three models I mentioned above
1
u/Confident_Ideal_5385 23m ago
Not easily, and not via prompting.
Your best bet is to just use gemma as a "renderer" for a plan assembled by Qwen.
1
u/DerTomsn 16h ago
Maybe you want to give this a try: https://www.reddit.com/r/LocalLLaMA/comments/1wv8uci/hemmingway1oq8emtp_up_to_320_toks_for_local/
17
u/Fluxing_Capacitor 15h ago
Creative writing and coding tasks are somewhat opposing goals, which is why you see STEM/Coding focused models perform worse at creative writing. Why are you trying to bolt creative writing into Qwen instead of using a model you already know can do the tasks better?