{"id":49,"date":"2026-07-29T18:58:50","date_gmt":"2026-07-29T18:58:50","guid":{"rendered":"https:\/\/bvandermaas.com\/?p=49"},"modified":"2026-07-29T19:15:32","modified_gmt":"2026-07-29T19:15:32","slug":"can-we-actually-rightsize-model-usage","status":"publish","type":"post","link":"https:\/\/bvandermaas.com\/?p=49","title":{"rendered":"Can we actually rightsize model usage?"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">In this new reality where we have the privilege to not only take apart cloud infrastructure and re-engineer it for efficiency, but also bite down on AI implementations, it is becoming increasingly clear that we need strategies for optimisation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">One of the most important handles to finetune a workload in the cloud is rightsizing a VM. That can be through classic resizing, moving to a container set-up, moving to an autoscaling solution, going serverless, \u2026 to be fair, that last one is not really rightsizing, but the concept is there: making sure you run the appropriate compute resources for your workload.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Now, consider the following:<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You are chatting with your favourite new LLM, asking it to fix a typo in an email, a script, &#8230; Under the hood, that little request just walked through the same reasoning stack that could solve a PhD level physics problem, or a <a href=\"https:\/\/www.scientificamerican.com\/article\/ai-just-solved-an-80-year-old-erdos-problem-and-mathematicians-are-amazed\/\">complex math problem<\/a>. Nothing was rightsized whatsoever. You just drove your Ferrari from one end of your driveway to another top open your mailbox, basically.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">That is the state of AI spend right now.<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Everyone is default routing every prompt to the biggest brain available, the same way we used to spin up an m5.4xlarge to run a task Lamba can perfectly handly. Slowly we are starting to fix that behavior in the cloud. For AI however, it almost feels like none of the lessons from Cloud are taken into account.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The instance family problem, but for intelligence<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Rightsizing a VM starts with one simple question: \u201cWhat does this workload actually need?\u201d Not what it <em>might<\/em> need on its worst day, not what feels safe, what it needs right now, based on evidence. Yes, <em>might<\/em> and <em>safe<\/em> should play a role, but if you do the cloud right you let the cloud handle deviating scenarios. (Read: elasticity)<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Model selection deserves the exact same diligence. Summarizing a support ticket does not need the same advanced model that is refactoring five interdependent services at 2am. But because switching models still feels like a decision, a manual toggle, most people just leave it on the expensive default and eat the bill. What probably also is a factor, is that there\u2019s still a desire for models to improve and a hope that new models are more and more reliable. While that\u2019s understandable, this is why you\u2019re probably bleeding money.Anyone who has ever fought a team out of oversized reserved instances knows this pattern by heart.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Enter the router<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Some products landed recently that made this analogy impossible to ignore, and they came from very different corners of tech.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Cursor<\/strong> shipped a Router that sits inside the IDE and classifies every single coding request before it ever touches a model. Simple, familiar edits get bounced to something cheap and fast. A gnarly multi-file refactor gets escalated to a frontier model. You pick a dial, Intelligence, Balance, or Cost, and the router steers along that frontier for you. Think of it as an autoscaling group, except instead of scaling instance count based on CPU load, it is scaling model tier based on task complexity. Same instinct, new axis.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Ramp<\/strong> took the harder, more forensic road. Their internal gateway does not just classify and forward, it learns. It watches latency and failure rates per model per hour, notices things like a discount tier quietly degrading during business hours, and adjusts routing in real time using statistical methods built for exactly this kind of decision under uncertainty.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That is a system running its own continuous audit of the fleet, the same forensic reflex FinOps people apply to spot pricing and reserved instance coverage, just pointed at tokens instead of vCPUs. Both approaches land in the same place: send the request to the cheapest model that still clears the bar. Everything else is implementation detail.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>There&#8217;s however still an important gap:<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Here is the part that should make every FinOps practitioner sit up. Rightsizing a VM has over a decade of tooling behind it. CPU and memory utilization graphs, sizing recommendations, guardrails, rollback plans, precedent and an established accepted culture. If you get it wrong, you know within a day and you fix it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Model rightsizing has none of that maturity yet. &#8216;Clears the bar is doing an enormous amount of unaudited work in every one of these router pitches. Whose bar. Measured how. <strong>Nobody is publishing the equivalent of a CloudWatch dashboard for reasoning quality<\/strong>, at least not yet.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Which is exactly the gap a forensic FinOps practice should be walking straight into. We already know how to ask uncomfortable questions about whether a resource matches the evidence of need. We just have not built the token era version of that discipline. The utilization metric of tomorrow is not CPU percent, it is something closer to task complexity per dollar, and right now almost nobody outside these router vendors is even trying to measure it independently.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><br>Rightsizing was never really about the VM. It was about refusing to pay for capability you are not using. So let&#8217;s apply that same logic to AI tasks, where we try and run only for what we actually need to achieve.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>In this new reality where we have the privilege to not only take apart cloud infrastructure and re-engineer it for efficiency, but also bite down on AI implementations, it is becoming increasingly clear that we need strategies for optimisation. One of the most important handles to finetune a workload in the cloud is rightsizing a [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":53,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[4],"tags":[],"class_list":["post-49","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-finops"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Can we actually rightsize model usage? - Cloud &amp; AI Efficiency Research<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/bvandermaas.com\/?p=49\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Can we actually rightsize model usage? - Cloud &amp; AI Efficiency Research\" \/>\n<meta property=\"og:description\" content=\"In this new reality where we have the privilege to not only take apart cloud infrastructure and re-engineer it for efficiency, but also bite down on AI implementations, it is becoming increasingly clear that we need strategies for optimisation. One of the most important handles to finetune a workload in the cloud is rightsizing a [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/bvandermaas.com\/?p=49\" \/>\n<meta property=\"og:site_name\" content=\"Cloud &amp; AI Efficiency Research\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-29T18:58:50+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-29T19:15:32+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/bvandermaas.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-29-at-21.14.28.png\" \/>\n\t<meta property=\"og:image:width\" content=\"914\" \/>\n\t<meta property=\"og:image:height\" content=\"414\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"benjaminvandermaas\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"benjaminvandermaas\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"4 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49\"},\"author\":{\"name\":\"benjaminvandermaas\",\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/#\\\/schema\\\/person\\\/ddf82a33603bc9f28078735ee615e8b9\"},\"headline\":\"Can we actually rightsize model usage?\",\"datePublished\":\"2026-07-29T18:58:50+00:00\",\"dateModified\":\"2026-07-29T19:15:32+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49\"},\"wordCount\":868,\"commentCount\":0,\"image\":{\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/bvandermaas.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/Screenshot-2026-07-29-at-21.14.28.png\",\"articleSection\":[\"FinOps\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/bvandermaas.com\\\/?p=49#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49\",\"url\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49\",\"name\":\"Can we actually rightsize model usage? - Cloud &amp; AI Efficiency Research\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/bvandermaas.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/Screenshot-2026-07-29-at-21.14.28.png\",\"datePublished\":\"2026-07-29T18:58:50+00:00\",\"dateModified\":\"2026-07-29T19:15:32+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/#\\\/schema\\\/person\\\/ddf82a33603bc9f28078735ee615e8b9\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/bvandermaas.com\\\/?p=49\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49#primaryimage\",\"url\":\"https:\\\/\\\/bvandermaas.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/Screenshot-2026-07-29-at-21.14.28.png\",\"contentUrl\":\"https:\\\/\\\/bvandermaas.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/Screenshot-2026-07-29-at-21.14.28.png\",\"width\":914,\"height\":414},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/?p=49#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/bvandermaas.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Can we actually rightsize model usage?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/#website\",\"url\":\"https:\\\/\\\/bvandermaas.com\\\/\",\"name\":\"FinOps Frontier\",\"description\":\"Investigating Cloud &amp; AI\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/bvandermaas.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/bvandermaas.com\\\/#\\\/schema\\\/person\\\/ddf82a33603bc9f28078735ee615e8b9\",\"name\":\"benjaminvandermaas\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/9ac4a26a67b13269d4b4a8fbe6fd17a3e3f2df75bf2b068fcd2a4cace0a0fbcf?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/9ac4a26a67b13269d4b4a8fbe6fd17a3e3f2df75bf2b068fcd2a4cace0a0fbcf?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/9ac4a26a67b13269d4b4a8fbe6fd17a3e3f2df75bf2b068fcd2a4cace0a0fbcf?s=96&d=mm&r=g\",\"caption\":\"benjaminvandermaas\"},\"sameAs\":[\"https:\\\/\\\/bvandermaas.com\"],\"url\":\"https:\\\/\\\/bvandermaas.com\\\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Can we actually rightsize model usage? - Cloud &amp; AI Efficiency Research","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/bvandermaas.com\/?p=49","og_locale":"en_US","og_type":"article","og_title":"Can we actually rightsize model usage? - Cloud &amp; AI Efficiency Research","og_description":"In this new reality where we have the privilege to not only take apart cloud infrastructure and re-engineer it for efficiency, but also bite down on AI implementations, it is becoming increasingly clear that we need strategies for optimisation. One of the most important handles to finetune a workload in the cloud is rightsizing a [&hellip;]","og_url":"https:\/\/bvandermaas.com\/?p=49","og_site_name":"Cloud &amp; AI Efficiency Research","article_published_time":"2026-07-29T18:58:50+00:00","article_modified_time":"2026-07-29T19:15:32+00:00","og_image":[{"width":914,"height":414,"url":"https:\/\/bvandermaas.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-29-at-21.14.28.png","type":"image\/png"}],"author":"benjaminvandermaas","twitter_card":"summary_large_image","twitter_misc":{"Written by":"benjaminvandermaas","Est. reading time":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/bvandermaas.com\/?p=49#article","isPartOf":{"@id":"https:\/\/bvandermaas.com\/?p=49"},"author":{"name":"benjaminvandermaas","@id":"https:\/\/bvandermaas.com\/#\/schema\/person\/ddf82a33603bc9f28078735ee615e8b9"},"headline":"Can we actually rightsize model usage?","datePublished":"2026-07-29T18:58:50+00:00","dateModified":"2026-07-29T19:15:32+00:00","mainEntityOfPage":{"@id":"https:\/\/bvandermaas.com\/?p=49"},"wordCount":868,"commentCount":0,"image":{"@id":"https:\/\/bvandermaas.com\/?p=49#primaryimage"},"thumbnailUrl":"https:\/\/bvandermaas.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-29-at-21.14.28.png","articleSection":["FinOps"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/bvandermaas.com\/?p=49#respond"]}]},{"@type":"WebPage","@id":"https:\/\/bvandermaas.com\/?p=49","url":"https:\/\/bvandermaas.com\/?p=49","name":"Can we actually rightsize model usage? - Cloud &amp; AI Efficiency Research","isPartOf":{"@id":"https:\/\/bvandermaas.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/bvandermaas.com\/?p=49#primaryimage"},"image":{"@id":"https:\/\/bvandermaas.com\/?p=49#primaryimage"},"thumbnailUrl":"https:\/\/bvandermaas.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-29-at-21.14.28.png","datePublished":"2026-07-29T18:58:50+00:00","dateModified":"2026-07-29T19:15:32+00:00","author":{"@id":"https:\/\/bvandermaas.com\/#\/schema\/person\/ddf82a33603bc9f28078735ee615e8b9"},"breadcrumb":{"@id":"https:\/\/bvandermaas.com\/?p=49#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/bvandermaas.com\/?p=49"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/bvandermaas.com\/?p=49#primaryimage","url":"https:\/\/bvandermaas.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-29-at-21.14.28.png","contentUrl":"https:\/\/bvandermaas.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-29-at-21.14.28.png","width":914,"height":414},{"@type":"BreadcrumbList","@id":"https:\/\/bvandermaas.com\/?p=49#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/bvandermaas.com\/"},{"@type":"ListItem","position":2,"name":"Can we actually rightsize model usage?"}]},{"@type":"WebSite","@id":"https:\/\/bvandermaas.com\/#website","url":"https:\/\/bvandermaas.com\/","name":"FinOps Frontier","description":"Investigating Cloud &amp; AI","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/bvandermaas.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/bvandermaas.com\/#\/schema\/person\/ddf82a33603bc9f28078735ee615e8b9","name":"benjaminvandermaas","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/9ac4a26a67b13269d4b4a8fbe6fd17a3e3f2df75bf2b068fcd2a4cace0a0fbcf?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/9ac4a26a67b13269d4b4a8fbe6fd17a3e3f2df75bf2b068fcd2a4cace0a0fbcf?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/9ac4a26a67b13269d4b4a8fbe6fd17a3e3f2df75bf2b068fcd2a4cace0a0fbcf?s=96&d=mm&r=g","caption":"benjaminvandermaas"},"sameAs":["https:\/\/bvandermaas.com"],"url":"https:\/\/bvandermaas.com\/?author=1"}]}},"brizy_media":[],"_links":{"self":[{"href":"https:\/\/bvandermaas.com\/index.php?rest_route=\/wp\/v2\/posts\/49","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/bvandermaas.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/bvandermaas.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/bvandermaas.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/bvandermaas.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=49"}],"version-history":[{"count":1,"href":"https:\/\/bvandermaas.com\/index.php?rest_route=\/wp\/v2\/posts\/49\/revisions"}],"predecessor-version":[{"id":50,"href":"https:\/\/bvandermaas.com\/index.php?rest_route=\/wp\/v2\/posts\/49\/revisions\/50"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/bvandermaas.com\/index.php?rest_route=\/wp\/v2\/media\/53"}],"wp:attachment":[{"href":"https:\/\/bvandermaas.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=49"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/bvandermaas.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=49"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/bvandermaas.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=49"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}