Local benchmark · manual recipe v1 · 13 September 2026

Reproduce a local AI test

Run the same short task on another PC. Keep the model, settings and evidence aligned before comparing results.

No browser scan, automatic download or execution. The recipe is a checklist, not an importable Cockpit script.

What the existing case actually shows

OutilsIA has two published historical observations on the same RTX 4080 SUPER 16 GB PC. They are useful examples of the evidence to record, not a controlled comparison.

ObservedHermes 3 8BQwen 3.8 27B
Date27 July 202619 August 2026
Modelhermes3:8bqwen3.8:27b, Q4_K_M
RuntimeOllama on WindowsOllama 0.32.14 on WSL2
Generation118.5 tokens/s9.4 tokens/s
GPU placement100% observed72% GPU / 28% CPU

Sources: public measurements (JSON), Qwen run detail (JSON) and French evidence register.

Different models, runtimes and settings prevent attributing the speed gap to GPU placement alone. These records do not contain everything needed for an exact historical replay. The recipe below starts a new series; it does not promise 118.5 tokens/s on your PC.

Your new test: Hermes 3 8B

Use hermes3:8b only if already installed and suitable for your machine. Otherwise stop here and choose a smaller model with Cockpit; that becomes a different test series. No large model download is required to participate.

Recipe: OIA-RECIPE-HERMES3-8B-20260913
Reference app: 0.1.2-beta, build 20260913142000. Public download currently: 0.1.2-beta / 20260913142000.

  1. Record the identity before startingNote app build, OS, Ollama version and Windows/WSL/Linux environment. Record the exact model digest and quantization from the installed model information. If unavailable, write “unknown”; do not call two results strictly comparable. A tag alone can change over time.
  2. Open the existing standard testScan in Cockpit, open « Tests », then « Préparer le test standard ». Check that « Modèle » is exactly hermes3:8b. This button may initially select another installed model. Do not click « Tester » if the intended model is missing.
  3. Keep the standard settingsUse the normal benchmark, not CPU retest, Arena or Autopilot. If an Autopilot profile is active, use « Restaurer » first. The standard Rust path uses context 2048, maximum 96 generated tokens, seed 42, temperature 0 and thinking disabled for Hermes. The app adds the short French-answer instruction shown below.
  4. Run once to warm up, then three recorded runsReview and confirm each benchmark. Run sequentially, promptly after the previous run, with other heavy work closed. Keep the same prompt, runtime and settings. The standard Hermes timeout is 45 seconds per run. A timeout is a failed run, not permission to retry indefinitely.
  5. Keep all three outcomesRecord generation count and duration, loading and prompt processing separately, and CPU/GPU placement if observed. Preserve failures. A median requires three successful comparable runs; keep individual results visible. This short test measures throughput, not general model quality.

Text to paste into “Prompt court”

Pourquoi la VRAM est importante pour un LLM local ?

Keep the French wording even in an English interface. Cockpit appends « Réponse finale uniquement, une phrase courte en français. ». Translating the task would change the experiment.

Download checklist JSON

The JSON records the protocol and missing identity fields; it contains no result and cannot execute commands. Cockpit 0.1.2 does not import this checklist automatically.

When is a comparison meaningful?

Match model digest, quantization, prompt, context, generation limit, app build, Ollama version and runtime environment. For a hardware comparison, change the PC deliberately and label that difference. For a before/after fix, keep the PC and change only the setting being investigated.

Generation speed is eval_count × 1000 / eval_duration_ms. Wall-clock response time includes more than generation. Do not substitute it for Ollama's generation duration. Missing counters stay missing; a screenshot or matching hash does not independently certify the origin of a result.

Contribute an external PC test

We are looking for Windows/Linux x64 PCs without a dedicated GPU, older NVIDIA GPUs, mid-range GPUs and shared-memory PCs. The current package is not for macOS or ARM64.

Keep a private copy first. Use Cockpit's « Contribution volontaire » preview if available, or prepare a redacted report. Remove names, local paths, IPs, tokens, prompts and raw model output. Share only hardware, software identities, settings, measured counters and limits.

Contact OutilsIA about a test · email: [email protected]. No automatic sending occurs. Publishing a case requires separate agreement and review; an unrun slot stays unmeasured.