In an empty working folder, build a small demonstration of model tiers with a pretend model: no network, no real model, no
real data. The pretend model gets a tier (small, medium, large) and a task, and answers right or wrong depending on how hard the
task is. A chooser starts every task on the tier listed for it in a plain file, runs a cheap check on the answer, moves up a
tier when the check rejects it, and counts the calls per tier. Keep the model and the check replaceable. Test it yourself:
show a run with every task starting too low, a run with the right starting tiers, and the check rejecting a wrong answer.
