On this page
Classify the work
Put tasks into three buckets. High-leverage work includes architecture, hard debugging, and risky refactors. Routine work includes tests, documentation, and narrow edits. Local-friendly work includes summarization, formatting, and private repetitive tasks.
Reserve premium capacity for the bucket where it changes the outcome. A strong model does not need to spend its best window renaming variables.
Add the usage signal
Before starting, look at remaining allowance and reset time for the top two candidate tools. If one has plenty of short-window capacity but a strained weekly budget, give it a bounded task. If a reset is close, delay the long run and use the gap for review.
For local models, replace remaining allowance with memory, context, queue state, and expected speed.
Keep your usage visible while you work. Get Super Notchy for $27 ↗
Keep the rule simple
A practical order is: choose the best-fit model, check whether the task fits its remaining window, then choose the next-best option if it does not. Do not optimize every prompt in real time.
A side-by-side view such as Super Notchy makes the check fast enough to become habit. Review results weekly and adjust the routing rule when one tool consistently delivers more value for a task category.