2025 multi-model work notes: documents, code, and usage limits
Revised September 7, 2026: this article preserves a personal workflow record from 2025. Unsupported model rankings and chip or supply-chain predictions have been removed. The earlier comparison also incorrectly grouped MCP and LangChain with hardware inference portability. These historical experiences do not describe current models or plans.
The division of work I used then
When I wrote these notes in 2025, I switched between GPT, Claude, Codex, and Cursor. GPT was my main tool for everyday discussion, Claude for organizing documents and breaking down architecture, and Codex or Cursor for implementation and refactoring. I did not control tasks, versions, or retries, so this routine cannot establish a model ranking.
I had high hopes for Cursor with Claude, but reached a usage limit after a few hours at the time. That affected my subscription decision. It does not mean everyone would reach the same limit at the same time: tasks, context, and the plan in effect also matter.
More documents mean more to maintain
Claude readily produced documents, so I regularly cleaned them up. The planning, architecture, and technical analysis I kept were useful, but not every output deserved ongoing maintenance. Each additional document could drift away from the code.
For a similar handoff, write down the next task, the decisions already settled, and the operation or artifact that will demonstrate completion. Include relevant files and constraints. This makes omissions easier to inspect than an unfiltered chat transcript. It is a suggested handoff method, not a measured efficiency result.
Keep a comparable task before changing models
When an answer disappoints, switching tools is tempting. But if the second prompt is clearer, the improvement cannot all be attributed to the model. Preserve the input, required files, and completion criteria before comparing results.
| Record | Purpose |
|---|---|
| Model, tool, and date | Identify version and environment differences |
| Input and completion criteria | Avoid comparing different tasks |
| Errors remaining after verification | Separate fluent output from usable work |
| Correction time and interruptions | Include manual rework, limits, and tool failures |
For a refactoring task, compilation is one check; replaying the original user operations checks whether behavior survives. For document organization, trace conclusions to sources and check that proposals have not become settled decisions. Keeping failures alongside successes makes the record more useful for the next choice.
Start with the work, not a platform
One stable task may need only one well-understood tool. Add another step when a concrete limit, capability gap, or handoff problem justifies it. Multiple models introduce context transfer and verification costs; more tools alone do not establish reliability.
The useful part of this historical record is to define the work and note where a tool helped or interrupted it. Plans and models change. Concrete handoff and verification records make those changes possible to assess.