We measure. You get the data.
Fit is computed. Performance is measured. Both are documented.
FitMyLLM answers one question: which model runs on which hardware, and how fast. The fit and sizing side is an engine, published and checkable. The speed side is a lab — we rent real machines, load real weights, time them, and release the method with the numbers.
If you sell a model, a GPU, or a machine that runs them, we can measure it and publish the result. If you are about to deploy one, we can measure it privately and tell you what it will cost. What follows is what each of those costs, and the rules the public work comes with.
Machines measured to date: NVIDIA GeForce RTX 4090, NVIDIA GeForce RTX 3090. Every report carries the llama.cpp build, the driver version, the date and the billed cost of producing it. Read them →
What you can buy
Commissioned public benchmark
from 3,000 CHFWe measure your model — or your machine — on hardware observed across our local-LLM audience, and publish it with the raw data at a permanent URL.
Starting scope: one model, one hardware configuration, two context lengths, and concurrency to an agreed ceiling. Larger matrices are quoted separately — a multi-hardware report typically runs 6,000–10,000 CHF, and a large comparative study is priced on the design.
Published under the editorial rules below.
Private sizing report
from 5,000 CHFThe same measurement stack pointed at your decision instead of at the public: which machine, how many, what it serves, cost per million tokens against your current API bill, and the break-even volume — including the case where you never reach it, which we will print as such.
From 5,000 CHF for a single workload and a limited set of candidate configurations. Multi-workload studies, or ones spanning hardware we do not already hold, are scoped and quoted individually.
Confidential. Never published without a separate written agreement.
Sponsor placement
500–1,500 CHF / monthA labelled, visually separated placement outside any scored list — the header of the GPU price tracker, the footer of Match results, or in-article. It never touches a ranking, a score or a recommendation.
Usually bought alongside a benchmark rather than on its own: a commissioned benchmark with three months of placement is 4,500 CHF, and a benchmark plus a year of placement is 10,000–15,000 CHF. Standalone: three concurrent placements maximum site-wide, three-month minimum term, invoiced quarterly, annual agreements on request. Category exclusivity priced separately. Monthly reporting of impressions and outbound clicks included, with the source and counting method stated in every report.
Labelled, and excluded from every ranking.
Dataset / API licence
price on requestOur joined fit, speed, quality and price dataset for local inference, as a feed rather than a page.
Coverage is still expanding — we quote this honestly against what exists on the day you ask, and will tell you when it does not yet cover your question.
Licensed, with the same provenance fields we publish.
What the deliverable actually looks like
Not a slide deck. Every engagement produces the same artifact the site already publishes, so you can read one in full before you commission anything — or walk the twelve sections one at a time, each shown with a real excerpt.
SAMPLE REPORTand the raw record behind it: /api/rig/rtx-4090-24gb
- Fit — weights on disk, measured peak VRAM, headroom at the stated context, and whether it loads at all.
- Single stream — decode and prefill tok/s with standard deviation, TTFT p50 and p95, load time.
- Under load — aggregate and per-stream throughput, TTFT p95 and TPOT p95 across concurrency, failed requests.
- Cost — per million tokens, single-stream and batched, from the machine’s billed hourly rate and the throughput we measured.
- Energy — board power sampled at the card during generation rather than taken from the TDP, as watts, kWh per million tokens and tokens per watt.
- Failures — every configuration that did not load, with the stage and the error. The published reports happen to have none, and say so: the attempted, succeeded and failed counts are printed either way.
- Provenance — pinned runtime build, driver, GPU count declared and detected, timestamps.
- The raw JSON — the structured record behind every figure, released with the report.
Decisions this answers
A benchmark is not the deliverable; a decision is. These are the questions the measurements above are built to settle, and the ones worth commissioning work for.
Whether your model fits, at what context, on the cards this audience actually reports — and what it does when it does not fit.
The quant, context and flags that minimise time-to-first-token without collapsing throughput, measured rather than assumed.
Cost per million tokens at real concurrency against a stated API price, including the case where the break-even volume is never reached.
The configurations that fail, the context at which loading stops, and how far performance falls once the cache is full.
The terms, which are the product
The scope is configurable. The editorial safeguards are not.
In this community a benchmark that reads as bought is worth less than no benchmark at all — to you as much as to us. Models, hardware, context lengths, concurrency, timing and price are all things we agree together. The four clauses below are what make the result citable, and they are the ones we do not move on.
Commissioned public benchmarks follow the publication rules below. Private sizing reports do not: they are confidential, they are not published, and nothing in them appears on this site without a separate written agreement. Anything received under NDA — unreleased weights, credentials, internal volumes, pricing — stays confidential in both kinds of engagement.
Every figure is reproducible from what we publish: the runtime build, the driver, the exact weights, the command. The structured dataset ships with the report at a permanent URL.
You see the results before publication and can respond in writing, published alongside. You can require correction of a factual error, or of a run executed differently from the agreed scope — we will re-run it. You cannot require a change to a correct number, or stop it going out.
If your model is slower than the alternative, or your machine does not fit what you hoped it would, that is in the public report. This clause is the entire reason anyone believes the good results.
Sponsored work is labelled as sponsored and excluded from every scored list, leaderboard and recommendation on the site. We do not sell ranking, position, or a score. There is no price for that.
Audience intelligence
Two numbers, and the second is the one that is hard to buy elsewhere. The first is reach: 27,235 visitors and 60,149 page views over 17 July – 16 August 2026, read from Vercel Web Analytics on 2026-08-16. Weekly visitors went from 2,841 to 6,782 across mid-May to mid-August 2026 — no month down.
The second is the mix of hardware reported by people who came here specifically to run a model locally: a distribution we have not seen published anywhere else, because it describes people actively evaluating local LLMs rather than the general gaming or cloud market. Reach you can buy in a dozen places. That, you cannot.
| WHERE THEY LAND | VISITORS | SHARE |
|---|---|---|
| GPU guides — /blog/gpu/… | 14,046 | 52% |
| GPU pages — /gpu/… | 5,846 | 21% |
| The model matcher — / | 5,613 | 21% |
| Model pages — /model/… | 4,181 | 15% |
Three quarters of it arrives on a page about one specific graphics card, which is a person deciding whether to buy that card. Top markets by visitors: United States 5,684, Germany 1,892, United Kingdom 982, France 775, Canada 675, Australia 653, Netherlands 555.
One caveat we would rather state than have you find: Singapore, China and Hong Kong together contribute roughly 7,200 visitors at 1.2–1.6 pages each, against 2.2–3.4 for the markets above. Some of that is automated. Read the engaged audience as about 20,000 a month and price against that, not against the headline.
| GPU | DETECTIONS | SHARE OF ELIGIBLE |
|---|---|---|
| Apple M1 (8GB) | 551 | 9.7% |
| NVIDIA GeForce RTX 4060 | 535 | 9.4% |
| NVIDIA GeForce RTX 3060 12 GB | 184 | 3.2% |
| Apple M4 (16GB) | 159 | 2.8% |
| AMD Radeon RX 9070 XT | 149 | 2.6% |
| Apple M1 Pro (16GB) | 146 | 2.6% |
| Apple M4 Pro (24GB) | 140 | 2.5% |
| NVIDIA Quadro P3200 Mobile | 123 | 2.2% |
| NVIDIA GeForce RTX 5070 Ti | 120 | 2.1% |
| NVIDIA GeForce RTX 5090 | 116 | 2.0% |
| Apple M3 Pro (18GB) | 108 | 1.9% |
| NVIDIA GeForce RTX 3070 | 93 | 1.6% |
| Total records | 10,079 |
| Mobile, integrated or unrecognised, excluded | -3,385 |
| Low-confidence GTX 980 records, excluded | -1,001 |
| Eligible GPU detections — the denominator | 5,693 |
- A detection is one record per browser session, written once when the home page reads the WebGL renderer string. No cookie, no account, no IP stored.
- Deduplicated within a session, and rate-limited to one per IP per hour. It is not deduplicated across sessions or devices, so these are sessions, not people — a returning visitor is counted again. Read it as a distribution, not as a headcount, and do not read the total as an audience size.
- Only renderer strings matching a GPU in our catalogue are recorded; anything unrecognised is discarded rather than guessed, which biases the set toward hardware we already know.
- The denominator for the share column is all eligible detections (5,693) — excluding phones and integrated graphics, and retaining Apple Silicon, which is an integrated GPU on paper and a common local-inference machine in practice.
- 1,001 detections reporting a GeForce GTX 980 are set aside: a 2014 card cannot outrank every current flagship, and the string matches the generic renderer returned by privacy-hardened browsers. Excluded from both the table and the denominator, present in the raw distribution.
- Collection gap: the write path failed silently from 16 to 26 July 2026 and recorded nothing. Everything before and after is intact; that window is missing and we are not going to interpolate it.
10,079 detections in total since 2026-03-21, across 249 distinct GPUs. The full distribution, and the traffic and behaviour figures, are in the media kit on request.
We run two analytics sources and they disagree, so we say it first rather than being asked. Vercel Analytics is cookieless and counts everyone; PostHog is consent-gated and will always report lower. We quote reach from the former and behaviour from the latter, and we will tell you which number came from where. A site whose entire pitch is published methodology does not get to be vague about its own.
How a commissioned benchmark runs
- We agree the scope in writing: which models, which hardware, which context lengths, which concurrency levels, and what would count as a negative result.
- We rent the machines, or you give us access to yours. Either way the provenance goes in the report.
- The harness runs unattended and writes a structured report — the same one behind every page under /rig.
- You get the draft and the raw data, and have ten working days to reply or to flag a factual error or an out-of-scope run.
- Public engagements: we publish at a permanent URL with your reply included, labelled sponsored, excluded from rankings. Private engagements: delivery, and nothing else.
Typical turnaround from signature to delivery is two to three weeks, most of which is measurement time and your reply window.
Tell us what you want measured and what you would consider a fair test. If the answer is already public, or the measurement would not be meaningful, we will say so rather than take the work.
fitmyllm@gmail.com