Find the right models for your tasks.

I previously talked about the hype, the lessons of past "revolutions," the ethics, and more recently, where things stand in the workforce. This time I want to talk about something more personal: the models themselves, and how I have actually been using them.

Over the past while, I have gone past a single assistant. I have used Claude, ChatGPT, Gemini, Grok, DeepSeek, Qwen, and Kimi, all on their free tiers, and I have run a couple of modified Ollama models myself, one in a container on Franky and another on an old Dell. The goal was never to crown a winner. It was to figure out what each one is actually good at.

The prompt is only half the story

The deeper I got into this, the clearer it became that the quality of a result depends less on raw capability and more on how a model interprets what you are asking for. The same prompt, sent to different models, can come back with wildly different results. That is the other half: which model you hand the job to, and sometimes how you phrase the job for that model.

A tale of three models

A good example: I asked Kimi, DeepSeek, and Grok to do the same thing, mimic the tone and the look of a particular web page in a document. The task was detailed and specific. I gave each of them a visual sample and a code sample of what to match, in style and layout.

All three tried to interpret the task from the prompt, and they still came back different. Kimi seemed to understand the meaning of the task. The result was almost identical in tone and layout, and the task itself was extremely well executed. Grok was closer than DeepSeek in look and feel, but I still needed a couple of follow-up prompts to get the result I wanted. DeepSeek felt like it did not quite feel the task. I tried the same kind of follow-up with it, and even after that it was not quite there.

This was one very specific task. I have tried other combinations, and the results shift. Each model has strengths the others do not, and the prompt sometimes has to be adjusted for the model in front of you.

One orchestrator, many specialists

That experience shifted how I think about using AI day to day. Instead of hunting for a single best model to do everything, the more effective setup is to have one model orchestrate and carry the bulk of the work, while other models step in for the specific things they do well, or the specific things you have tuned them to do. Less a single assistant, more a small team, each member doing what they are actually good at.

This is only the high-level view. There is more to unpack, especially around the modified Ollama models on Franky and the old Dell, but that is a story for another post.

So that's all for now.