Add OpenAI (GPT-5.6 Sol) as a second vision provider; gate server settings to a named admin

vision.py now dispatches per-model to _call_anthropic or _call_openai --
same prompt, same schema, same cardimage.py measurements either way, only
the request/response shape differs. Confirmed the existing GRADING_SCHEMA
already satisfies OpenAI's strict-mode requirement (every property listed
in required, additionalProperties:false at every level) with no changes.

Settings gained a second axis: which server key applies now depends on the
selected model's provider, and friends' personal keys are stored per
provider (with a one-time migration from the old single-key localStorage
slot) since a Claude key and an OpenAI key aren't interchangeable.

CARD_GRADER_ADMIN_USER names one username (read from the proxy's forwarded
basic-auth header) who alone may write server settings; everyone else keeps
the same read-only view CARD_GRADER_LOCK used to give everyone, while still
being able to set their own personal key. Deployed here as ninja_hippo.
CARD_GRADER_LOCK remains the fallback when no admin is named.
This commit is contained in:
Barely Removable 2026-08-23 07:39:03 -07:00
parent 9d5c68b093
commit 034761e145
7 changed files with 351 additions and 122 deletions

View file

@ -6,7 +6,7 @@ FROM python:3.12-slim
ENV PYTHONUNBUFFERED=1 ENV PYTHONUNBUFFERED=1
# Only optional deps (see README) — the app itself is stdlib only. # Only optional deps (see README) — the app itself is stdlib only.
RUN pip install --no-cache-dir anthropic pillow RUN pip install --no-cache-dir anthropic openai pillow
WORKDIR /app WORKDIR /app
COPY app.py cardimage.py store.py vision.py ./ COPY app.py cardimage.py store.py vision.py ./

View file

@ -101,41 +101,51 @@ graded — so regrading a card keeps its photos for another full window.
### Optional packages ### Optional packages
Two packages unlock real functionality. Each is checked for at runtime, so Three packages unlock real functionality. Each is checked for at runtime, so
the app runs without them — you just lose that capability. the app runs without them — you just lose that capability.
```bash ```bash
python3 -m pip install --user anthropic # required — this is what does the grading python3 -m pip install --user anthropic # Claude models — Sonnet 5, Haiku 4.5, Opus 5
python3 -m pip install --user openai # OpenAI models — GPT-5.6 Sol
python3 -m pip install --user pillow # the corner/edge/surface measurement pipeline python3 -m pip install --user pillow # the corner/edge/surface measurement pipeline
``` ```
Without `anthropic`, grading simply doesn't work — add a key in Settings once You only need whichever provider package matches the model you actually use
it's installed. Without `pillow`, grading still runs, but on the full photo — install both if you want the option to switch. Without `pillow`, grading
alone: no measured centering, no measured edge whitening, no magnified still runs, but on the full photo alone: no measured centering, no measured
close-ups. Corners and edges will come back `cannot_assess` far more often, edge whitening, no magnified close-ups. Corners and edges will come back
because a plain full-card photo genuinely doesn't show enough detail in those `cannot_assess` far more often, because a plain full-card photo genuinely
regions for a model to judge them from. Install `pillow` — it's the doesn't show enough detail in those regions for a model to judge them from.
difference between an answer and a shrug. Install `pillow` — it's the difference between an answer and a shrug.
### API key ### API keys
Get one at [console.anthropic.com](https://console.anthropic.com), then paste Get an Anthropic key at [console.anthropic.com](https://console.anthropic.com)
it into Settings in the app. Grading costs a few cents per card (Sonnet 5, and/or an OpenAI key at [platform.openai.com](https://platform.openai.com),
~$0.04-0.06 depending on how many photos you send) — nothing is charged until then paste whichever you use into Settings in the app — each provider has its
you click "Estimate the grade." own field, since the keys aren't interchangeable and you may want to hold
both. Grading costs a few cents per card either way (Sonnet 5 and GPT-5.6 Sol
are both currently ~$0.04-0.06 depending on how many photos you send) —
nothing is charged until you click "Estimate the grade."
## Why Sonnet 5, not a cheaper model ## Why Sonnet 5, not a cheaper model
Settings lets you switch to Haiku 4.5 (about a third of the cost) or Opus 5 Settings lets you switch to Haiku 4.5 (about a third of the cost), Opus 5
(more expensive, marginally more careful). Sonnet is the default for a (more expensive, marginally more careful), or GPT-5.6 Sol (OpenAI, similar
tested reason, not a hunch: run head-to-head against Haiku on the same price to Sonnet). Sonnet is the default for a tested reason, not a hunch: run
photos, Haiku inverted a PSA centering-tolerance comparison — read 59/41 as head-to-head against Haiku on the same photos, Haiku inverted a PSA
*exceeding* the tolerance for a 9, when 59/41 is actually well inside it — centering-tolerance comparison — read 59/41 as *exceeding* the tolerance for
and the error alone dragged its estimate three grade levels below Sonnet's. a 9, when 59/41 is actually well inside it — and the error alone dragged its
Haiku also lacks extended thinking entirely, which matters for exactly the estimate three grade levels below Sonnet's. Haiku also lacks extended
judgment calls this task is full of: is this a scratch or a print texture, thinking entirely, which matters for exactly the judgment calls this task is
a reflection or real edge wear, a print line or a crease. If cost matters full of: is this a scratch or a print texture, a reflection or real edge
enough to switch, that's a real, deliberate tradeoff — not a free lunch. wear, a print line or a crease. If cost matters enough to switch, that's a
real, deliberate tradeoff — not a free lunch.
GPT-5.6 Sol hasn't been run through the same head-to-head testing — it's
wired up and usable, but nobody has checked it against real graded cards the
way Haiku was checked here. Treat its results with the scrutiny that implies
until someone does that comparison.
## The workflow ## The workflow
@ -224,12 +234,14 @@ includes a `Dockerfile`, `docker-compose.yml`, and an nginx config
-d card-grader.yourdomain.com -d card-grader.yourdomain.com
docker restart nginx docker restart nginx
``` ```
5. **Set your Anthropic API key** by opening `https://card-grader.yourdomain.com` 5. **Set your API key(s)** by opening `https://card-grader.yourdomain.com`
yourself first, going to Settings, and pasting it in — this only works yourself first, going to Settings, and pasting them in. If you're using
*before* `CARD_GRADER_LOCK=1` takes effect, so either set the key first `CARD_GRADER_ADMIN_USER`, this just works — settings stay writable for
and then bring the lock up, or temporarily comment out that line in that one username indefinitely. If you're using the blanket
`docker-compose.yml`, restart, set the key, uncomment it, and restart `CARD_GRADER_LOCK=1` instead, that only works *before* the lock takes
again. effect, so either set the key first and then bring the lock up, or
temporarily comment out that line in `docker-compose.yml`, restart, set
the key, uncomment it, and restart again.
If Unraid already has an nginx-based reverse proxy set up (Nginx Proxy If Unraid already has an nginx-based reverse proxy set up (Nginx Proxy
Manager, SWAG, etc.), skip step 4 and just add a proxy host there pointing Manager, SWAG, etc.), skip step 4 and just add a proxy host there pointing
@ -268,11 +280,25 @@ key is being used**. Two things in the app deal with this:
CARD_GRADER_LOCK=1 python3 app.py CARD_GRADER_LOCK=1 python3 app.py
``` ```
which refuses all settings writes. Without it, any visitor could overwrite which refuses all settings writes from everyone. Without it, any visitor
your stored API key or switch the model to the most expensive one. Grading could overwrite your stored API key or switch the model to the most
still works normally; only the settings write is blocked. expensive one. Grading still works normally; only the settings write is
blocked.
Your stored API key is never sent to any browser regardless — the settings If your reverse proxy does per-user basic auth (as opposed to one shared
password for everyone), you can name a single admin instead of locking
settings out entirely:
```bash
CARD_GRADER_ADMIN_USER=yourname python3 app.py
```
Only that username may write settings; everyone else gets the same
read-only view `CARD_GRADER_LOCK` gives everyone. Either way, everyone can
still set their *own* personal key (point 1 above) — that never touches
server settings at all.
Your stored API keys are never sent to any browser regardless — the settings
endpoint returns only a boolean saying whether one is configured. endpoint returns only a boolean saying whether one is configured.
Everyone shares one `grades.db`, so the History list is communal. That's Everyone shares one `grades.db`, so the History list is communal. That's

53
app.py
View file

@ -76,6 +76,20 @@ TEMPLATED_STATIC = {"index.html", "manifest.json", "sw.js", "app.js"}
# normally; only the settings write is refused. # normally; only the settings write is refused.
SETTINGS_LOCKED = os.environ.get("CARD_GRADER_LOCK", "").strip() not in ("", "0") SETTINGS_LOCKED = os.environ.get("CARD_GRADER_LOCK", "").strip() not in ("", "0")
# When set, only this ONE username (from the reverse proxy's basic auth, see
# Handler._username) may write server settings — everyone else sees the same
# read-only view CARD_GRADER_LOCK produces, regardless of CARD_GRADER_LOCK's
# own value. This is strictly narrower than the blanket lock: it names one
# person rather than locking out or opening up to everyone at once. Blank
# (the default) falls back to CARD_GRADER_LOCK's all-or-nothing behaviour.
ADMIN_USERNAME = os.environ.get("CARD_GRADER_ADMIN_USER", "").strip()
def _settings_writable(username):
if ADMIN_USERNAME:
return username == ADMIN_USERNAME
return not SETTINGS_LOCKED
CONTENT_TYPES = { CONTENT_TYPES = {
".html": "text/html; charset=utf-8", ".html": "text/html; charset=utf-8",
".css": "text/css; charset=utf-8", ".css": "text/css; charset=utf-8",
@ -119,21 +133,27 @@ def _decode_uploaded_images(raw_images):
return images, None return images, None
def _public_settings(): def _public_settings(username):
"""Settings safe to hand to a browser. """Settings safe to hand to a browser.
The stored API key never leaves the server. Once this is reachable by Stored API keys never leave the server. Once this is reachable by anyone
anyone but you which is the whole point of putting it behind a tunnel but you which is the whole point of putting it behind a tunnel for
for friends returning the raw key would hand every visitor the ability friends returning the raw key would hand every visitor the ability to
to spend your Anthropic credits anywhere they like. They get a boolean spend your credits anywhere they like. They get a boolean per provider
saying whether one is configured, which is all the UI needs. saying whether one is configured, which is all the UI needs.
""" """
s = store.get_settings() s = store.get_settings()
writable = _settings_writable(username)
return { return {
"vision_model": s.get("vision_model"), "vision_model": s.get("vision_model"),
"vision_effort": s.get("vision_effort"), "vision_effort": s.get("vision_effort"),
"server_key_configured": bool(s.get("anthropic_api_key")), "anthropic_key_configured": bool(s.get("anthropic_api_key")),
"settings_locked": SETTINGS_LOCKED, "openai_key_configured": bool(s.get("openai_api_key")),
"is_admin": writable,
# Kept for the frontend's existing "locked" UI treatment — now means
# "not writable by YOU", whatever the reason, rather than a single
# global flag.
"settings_locked": not writable,
} }
@ -308,7 +328,7 @@ class Handler(BaseHTTPRequestHandler):
if route.startswith("/static/"): if route.startswith("/static/"):
return self._static(route[len("/static/"):]) return self._static(route[len("/static/"):])
if route == "/api/settings": if route == "/api/settings":
return self._json(_public_settings()) return self._json(_public_settings(self._username()))
if route == "/api/vision-models": if route == "/api/vision-models":
return self._json(vision.price_guide()) return self._json(vision.price_guide())
if route == "/api/usage": if route == "/api/usage":
@ -337,12 +357,13 @@ class Handler(BaseHTTPRequestHandler):
if body is None: if body is None:
return # _body already sent 413 and closed return # _body already sent 413 and closed
if route == "/api/settings": if route == "/api/settings":
if SETTINGS_LOCKED: username = self._username()
if not _settings_writable(username):
return self._error( return self._error(
"Settings are locked on this server. Use your own API " "Settings are locked on this server. Use your own API "
"key in this browser instead.", 403) "key in this browser instead.", 403)
store.save_settings(body) store.save_settings(body)
return self._json(_public_settings()) return self._json(_public_settings(username))
if route == "/api/grade": if route == "/api/grade":
return self._grade_card(body) return self._grade_card(body)
if route.startswith("/api/history/") and route.endswith("/regrade"): if route.startswith("/api/history/") and route.endswith("/regrade"):
@ -399,11 +420,17 @@ class Handler(BaseHTTPRequestHandler):
wording. wording.
""" """
model = body.get("model") if body.get("model") in vision.MODELS else settings.get("vision_model") model = body.get("model") if body.get("model") in vision.MODELS else settings.get("vision_model")
# Which stored server key applies depends on which model this grade
# actually runs on — Sonnet needs the Anthropic key, GPT-5.6 Sol
# needs the OpenAI one, and they are never interchangeable.
provider = vision.MODELS.get(model, {}).get("provider", "anthropic")
server_key = settings.get("openai_api_key" if provider == "openai" else "anthropic_api_key")
# A key sent with the request wins over the server's own. That is what # A key sent with the request wins over the server's own. That is what
# makes this shareable: hand a friend the URL and they can bring their # makes this shareable: hand a friend the URL and they can bring their
# own Anthropic key rather than spending yours. Never stored — it # own key (matching whichever provider `model` needs) rather than
# lives in their browser and is used for this one call. # spending yours. Never stored — it lives in their browser and is
# used for this one call.
caller_key = (body.get("api_key") or "").strip() or None caller_key = (body.get("api_key") or "").strip() or None
print("[grade] calling vision model={} on {} image(s){}".format( print("[grade] calling vision model={} on {} image(s){}".format(
@ -411,7 +438,7 @@ class Handler(BaseHTTPRequestHandler):
t0 = time.time() t0 = time.time()
result = vision.grade_card( result = vision.grade_card(
images, images,
caller_key or settings.get("anthropic_api_key") or None, caller_key or server_key or None,
model=model, model=model,
effort=settings.get("vision_effort"), effort=settings.get("vision_effort"),
) )

View file

@ -27,10 +27,18 @@ services:
volumes: volumes:
- ./data:/data - ./data:/data
environment: environment:
# Refuses settings writes from anyone but the host, so a visitor # Only this ONE username (from NPM's basic auth, forwarded upstream)
# can't overwrite the API key or switch to a pricier model. Set to 0 # may write server settings — model choice, both provider API keys.
# temporarily (on the live server only, not here) while setting the # Everyone else gets the same read-only view CARD_GRADER_LOCK used to
# API key via the app's own Settings screen, then back to 1. # give everyone; they can still set their OWN personal key, which
# never touches server settings at all. Set with the app's Settings
# screen unlocked for ninja_hippo only, so paste keys in there, not
# here.
- CARD_GRADER_ADMIN_USER=ninja_hippo
# CARD_GRADER_LOCK is now redundant with CARD_GRADER_ADMIN_USER set
# (the admin check takes priority) — kept only as the fallback for
# anyone who unsets the admin var and wants the old all-or-nothing
# behaviour back.
- CARD_GRADER_LOCK=1 - CARD_GRADER_LOCK=1
# This app is reverse-proxied at hippofam.com/cards, not the domain # This app is reverse-proxied at hippofam.com/cards, not the domain
# root — see app.py's BASE_PATH handling. # root — see app.py's BASE_PATH handling.

View file

@ -295,7 +295,10 @@ async function runGrade() {
images: pickedToPayload(gradeState.files), images: pickedToPayload(gradeState.files),
label: ($('#grade-label').value || '').trim() || null, label: ($('#grade-label').value || '').trim() || null,
}; };
const key = myApiKey(); // A fresh grade always runs on the server's configured default model
// (there's no per-grade model picker), so that's whose provider the
// personal key needs to match.
const key = myApiKeyForModel(state.settings.vision_model);
if (key) body.api_key = key; if (key) body.api_key = key;
gradeState.result = await api('/api/grade', { method: 'POST', body }); gradeState.result = await api('/api/grade', { method: 'POST', body });
status.hidden = true; status.hidden = true;
@ -606,7 +609,7 @@ async function openSettings() {
`<option value="${id}" ${s.vision_model === id ? 'selected' : ''}> `<option value="${id}" ${s.vision_model === id ? 'selected' : ''}>
${esc(info.label)} ~$${info.per_grade.toFixed(3)}/grade</option>`).join(''); ${esc(info.label)} ~$${info.per_grade.toFixed(3)}/grade</option>`).join('');
const locked = s.settings_locked; const isAdmin = s.is_admin;
$('#settings-modal').innerHTML = ` $('#settings-modal').innerHTML = `
<div class="panel-head"> <div class="panel-head">
@ -615,39 +618,56 @@ async function openSettings() {
</div> </div>
<div class="section"> <div class="section">
<h3>Your API key</h3> <h3>Your API keys</h3>
<div class="fields"> <div class="fields">
<div class="field wide"> <div class="field wide">
<label>Anthropic API key (this browser only)</label> <label>Anthropic key (this browser only)</label>
<input class="input" id="set-my-key" type="password" value="${esc(myApiKey() || '')}" <input class="input" id="set-my-key-anthropic" type="password"
placeholder="sk-ant-..."> value="${esc(myApiKey('anthropic') || '')}" placeholder="sk-ant-...">
<span class="suffix">Stored only in this browser and sent with your own grading <span class="suffix">Used when the selected model is a Claude model. Stored only in
requests never saved on the server. Get one at console.anthropic.com. this browser, never saved on the server. Get one at console.anthropic.com.</span>
${s.server_key_configured </div>
? 'This server already has a key set up, so you can leave this blank and use that instead — but then the owner pays for your grades.' <div class="field wide">
: 'This server has no key of its own, so you need one here to grade anything.'}</span> <label>OpenAI key (this browser only)</label>
<input class="input" id="set-my-key-openai" type="password"
value="${esc(myApiKey('openai') || '')}" placeholder="sk-...">
<span class="suffix">Used when the selected model is GPT-5.6 Sol. Same deal this
browser only. Get one at platform.openai.com.</span>
</div>
<div class="field wide">
<span class="suffix">${
(s.anthropic_key_configured || s.openai_key_configured)
? 'This server already has a key set up for at least one provider, so you can leave the matching field above blank and use that instead — but then the owner pays for your grades.'
: 'This server has no key of its own yet, so you need one above to grade anything.'
}</span>
</div> </div>
</div> </div>
</div> </div>
<div class="section"> <div class="section">
<h3>Server settings${locked ? ' <span class="pill pill-raw">locked</span>' : ''}</h3> <h3>Server settings${isAdmin ? '' : ' <span class="pill pill-raw">locked</span>'}</h3>
${locked ? `<div class="note">This server is shared, so its settings are read-only. ${!isAdmin ? `<div class="note">Only the admin can change server settings on this
Use your own key above.</div>` : ` instance. Use your own key above.</div>` : `
<div class="fields"> <div class="fields">
<div class="field wide"> <div class="field wide">
<label>Server API key (used when a visitor has none)</label> <label>Server Anthropic key (used when a visitor has none, and the model is Claude)</label>
<input class="input" id="set-api-key" type="password" <input class="input" id="set-api-key-anthropic" type="password"
placeholder="${s.server_key_configured ? '•••••••• already set — type to replace' : 'sk-ant-...'}"> placeholder="${s.anthropic_key_configured ? '•••••••• already set — type to replace' : 'sk-ant-...'}">
<span class="suffix">Never sent back to the browser once saved. Leave blank to keep </div>
the current one.</span> <div class="field wide">
<label>Server OpenAI key (used when a visitor has none, and the model is GPT-5.6 Sol)</label>
<input class="input" id="set-api-key-openai" type="password"
placeholder="${s.openai_key_configured ? '•••••••• already set — type to replace' : 'sk-...'}">
<span class="suffix">Neither key is ever sent back to a browser once saved. Leave a
field blank to keep whatever's already stored for it.</span>
</div> </div>
<div class="field wide"> <div class="field wide">
<label>Vision model</label> <label>Vision model</label>
<select id="set-model">${modelOptions}</select> <select id="set-model">${modelOptions}</select>
<span class="suffix">Sonnet 5 is the default for a tested reason run head-to-head <span class="suffix">Sonnet 5 is the default for a tested reason run head-to-head
against Haiku on the same cards, Haiku misread a PSA centering tolerance and landed against Haiku on the same cards, Haiku misread a PSA centering tolerance and landed
three grades off. Drop to Haiku only if cost matters more than accuracy to you.</span> three grades off. GPT-5.6 Sol hasn't been run against real cards here yet, so treat
its results with more scrutiny until it has.</span>
</div> </div>
</div>`} </div>`}
</div> </div>
@ -660,33 +680,60 @@ async function openSettings() {
} }
/* A personal key lives in localStorage, never on the server that's what /* A personal key lives in localStorage, never on the server that's what
lets someone use a shared instance without spending the owner's credits. */ lets someone use a shared instance without spending the owner's credits.
function myApiKey() { Keyed per provider since a Claude key and an OpenAI key aren't
try { return localStorage.getItem('cardgrader_api_key') || ''; } catch (_) { return ''; } interchangeable and someone may reasonably hold both. */
} function myApiKey(provider) {
function setMyApiKey(value) {
try { try {
if (value) localStorage.setItem('cardgrader_api_key', value); if (provider === 'anthropic') {
else localStorage.removeItem('cardgrader_api_key'); // One-time migration: friends who set a key before OpenAI support
// existed had it under the old unprefixed name. Move it once rather
// than losing it.
const legacy = localStorage.getItem('cardgrader_api_key');
if (legacy && !localStorage.getItem('cardgrader_api_key_anthropic')) {
localStorage.setItem('cardgrader_api_key_anthropic', legacy);
localStorage.removeItem('cardgrader_api_key');
}
}
return localStorage.getItem(`cardgrader_api_key_${provider}`) || '';
} catch (_) { return ''; }
}
function setMyApiKey(provider, value) {
try {
const key = `cardgrader_api_key_${provider}`;
if (value) localStorage.setItem(key, value);
else localStorage.removeItem(key);
} catch (_) { /* private browsing — the field just won't persist */ } } catch (_) { /* private browsing — the field just won't persist */ }
} }
// Which of the two personal keys applies to whatever model is actually
// going to run — the currently configured default, unless a specific grade
// requests a different one (nothing does yet, but the lookup is already
// provider-aware for when it does).
function myApiKeyForModel(modelId) {
const provider = (state.modelGuide[modelId] || {}).provider || 'anthropic';
return myApiKey(provider);
}
function closeSettings() { function closeSettings() {
$('#settings-modal').hidden = true; $('#settings-modal').hidden = true;
$('#modal-scrim').hidden = true; $('#modal-scrim').hidden = true;
} }
async function saveSettings() { async function saveSettings() {
setMyApiKey($('#set-my-key').value.trim()); setMyApiKey('anthropic', $('#set-my-key-anthropic').value.trim());
setMyApiKey('openai', $('#set-my-key-openai').value.trim());
// Only push server settings when this instance allows it, and only send a // Only push server settings when this instance allows it, and only send a
// key when one was actually typed — an empty box means "leave it alone", // key when one was actually typed — an empty box means "leave it alone",
// not "erase it". // not "erase it".
const serverKeyField = $('#set-api-key'); const modelField = $('#set-model');
if (serverKeyField) { if (modelField) {
const payload = { vision_model: $('#set-model').value }; const payload = { vision_model: modelField.value };
const typed = serverKeyField.value.trim(); const typedAnthropic = $('#set-api-key-anthropic').value.trim();
if (typed) payload.anthropic_api_key = typed; const typedOpenai = $('#set-api-key-openai').value.trim();
if (typedAnthropic) payload.anthropic_api_key = typedAnthropic;
if (typedOpenai) payload.openai_api_key = typedOpenai;
state.settings = await api('/api/settings', { method: 'POST', body: payload }); state.settings = await api('/api/settings', { method: 'POST', body: payload });
} }
closeSettings(); closeSettings();
@ -717,8 +764,12 @@ async function load() {
updateModelCostHint(); updateModelCostHint();
await loadHistory(); await loadHistory();
loadUsage().catch(() => {}); loadUsage().catch(() => {});
if (!state.settings.server_key_configured && !myApiKey()) { const activeProvider = (state.modelGuide[state.settings.vision_model] || {}).provider || 'anthropic';
banner('Add your Anthropic API key in Settings before grading a card.'); const serverHasKey = activeProvider === 'openai'
? state.settings.openai_key_configured : state.settings.anthropic_key_configured;
if (!serverHasKey && !myApiKey(activeProvider)) {
banner(`Add your ${activeProvider === 'openai' ? 'OpenAI' : 'Anthropic'} API key in `
+ 'Settings before grading a card.');
} }
} }

View file

@ -77,6 +77,7 @@ CREATE INDEX IF NOT EXISTS idx_events_created ON grade_events(created_at DESC);
DEFAULT_SETTINGS = { DEFAULT_SETTINGS = {
"anthropic_api_key": "", "anthropic_api_key": "",
"openai_api_key": "",
"vision_model": "claude-sonnet-5", "vision_model": "claude-sonnet-5",
"vision_effort": "low", "vision_effort": "low",
} }

186
vision.py
View file

@ -1,4 +1,8 @@
"""Estimate a trading card's PSA grade from photographs, using Claude's vision. """Estimate a trading card's PSA grade from photographs, using a vision model.
Supports Anthropic (Claude) and OpenAI as interchangeable providers pick
one per grade via MODELS below; the prompt, schema and cardimage.py
measurements are identical either way, only _call_anthropic/_call_openai
differ.
This is the one part of the app that costs money per use, and the one part This is the one part of the app that costs money per use, and the one part
that can be confidently wrong a model eyeballing a photo can miss real wear that can be confidently wrong a model eyeballing a photo can miss real wear
@ -25,6 +29,11 @@ try:
except ImportError: # keeps the rest of the app importable without the SDK except ImportError: # keeps the rest of the app importable without the SDK
anthropic = None anthropic = None
try:
import openai
except ImportError:
openai = None
# Per-model capabilities. These differ in ways that are 400 errors, not # Per-model capabilities. These differ in ways that are 400 errors, not
# preferences, so the request is built from this table rather than assuming a # preferences, so the request is built from this table rather than assuming a
# single shape: # single shape:
@ -38,6 +47,7 @@ except ImportError: # keeps the rest of the app importable without the SDK
MODELS = { MODELS = {
"claude-haiku-4-5": { "claude-haiku-4-5": {
"label": "Haiku 4.5 — cheapest", "label": "Haiku 4.5 — cheapest",
"provider": "anthropic",
"effort": False, # sending output_config.effort is a 400 "effort": False, # sending output_config.effort is a 400
"adaptive_thinking": False, "adaptive_thinking": False,
"fallbacks": False, "fallbacks": False,
@ -46,6 +56,7 @@ MODELS = {
}, },
"claude-sonnet-5": { "claude-sonnet-5": {
"label": "Sonnet 5 — balanced", "label": "Sonnet 5 — balanced",
"provider": "anthropic",
"effort": True, "effort": True,
"adaptive_thinking": True, "adaptive_thinking": True,
"fallbacks": False, "fallbacks": False,
@ -58,12 +69,29 @@ MODELS = {
}, },
"claude-opus-5": { "claude-opus-5": {
"label": "Opus 5 — most accurate", "label": "Opus 5 — most accurate",
"provider": "anthropic",
"effort": True, "effort": True,
"adaptive_thinking": True, "adaptive_thinking": True,
"fallbacks": True, "fallbacks": True,
"in_per_mtok": 5.00, "out_per_mtok": 25.00, "in_per_mtok": 5.00, "out_per_mtok": 25.00,
"approx_image_tokens": 4800, "approx_image_tokens": 4800,
}, },
"gpt-5.6-sol": {
"label": "GPT-5.6 Sol (OpenAI) — balanced",
"provider": "openai",
# No effort/thinking control wired up for this provider yet — every
# call runs at whatever this model's default reasoning depth is.
"effort": False,
"adaptive_thinking": False,
"fallbacks": False,
"in_per_mtok": 2.00, "out_per_mtok": 10.00,
# A rough estimate, unlike the Anthropic figures (which were true'd
# up against real usage — see the README's grading-cost note). Only
# affects the ADVERTISED per-grade estimate in Settings; the actual
# billed cost always comes from the real usage this API call
# reports, never from this number.
"approx_image_tokens": 1500,
},
} }
# Grading rewards the extra reasoning a thinking-capable model does — telling # Grading rewards the extra reasoning a thinking-capable model does — telling
@ -91,12 +119,14 @@ class VisionError(Exception):
def available(): def available():
"""Is the feature usable right now?""" """Is the feature usable right now, on at least one provider?"""
return anthropic is not None and bool(_api_key()) return ((anthropic is not None and bool(_env_api_key("anthropic")))
or (openai is not None and bool(_env_api_key("openai"))))
def _api_key(): def _env_api_key(provider):
return os.environ.get("ANTHROPIC_API_KEY", "").strip() var = "ANTHROPIC_API_KEY" if provider == "anthropic" else "OPENAI_API_KEY"
return os.environ.get(var, "").strip()
def media_type(filename, image_bytes=None): def media_type(filename, image_bytes=None):
@ -115,35 +145,12 @@ def media_type(filename, image_bytes=None):
return SUPPORTED_MEDIA.get(ext) return SUPPORTED_MEDIA.get(ext)
def _call_vision(images, system, schema, prompt, api_key=None, model=None, def _build_content(images, labels, image_block):
effort=None, max_tokens=None, labels=None): """Shared across providers: captions + encoded images, in request order.
"""Shared plumbing for every vision call: auth, request shape per model
capability, and error/refusal handling. Returns (parsed_json, usage_dict).
Raises VisionError with a human-readable message on any failure callers `image_block(mime, b64)` returns the provider-specific dict for one
surface it in the UI rather than half-committing anything. image the two SDKs disagree on that shape, nothing else here differs.
""" """
if anthropic is None:
raise VisionError(
"The anthropic package isn't installed. Run: "
"python3 -m pip install --user anthropic"
)
if not images:
raise VisionError("No images to analyze.")
key = (api_key or _api_key()) or None
if not key:
raise VisionError(
"No Anthropic API key set. Add one in Settings, or export "
"ANTHROPIC_API_KEY before starting the app."
)
model = model if model in MODELS else DEFAULT_MODEL
caps = MODELS[model]
effort = effort or DEFAULT_EFFORT
client = anthropic.Anthropic(api_key=key)
content = [] content = []
for index, (image_bytes, filename) in enumerate(images): for index, (image_bytes, filename) in enumerate(images):
mime = media_type(filename, image_bytes) mime = media_type(filename, image_bytes)
@ -159,13 +166,58 @@ def _call_vision(images, system, schema, prompt, api_key=None, model=None,
if labels and index < len(labels) and labels[index]: if labels and index < len(labels) and labels[index]:
content.append({"type": "text", "text": labels[index]}) content.append({"type": "text", "text": labels[index]})
encoded = base64.standard_b64encode(image_bytes).decode("utf-8") encoded = base64.standard_b64encode(image_bytes).decode("utf-8")
content.append({"type": "image", content.append(image_block(mime, encoded))
"source": {"type": "base64", "media_type": mime, "data": encoded}}) return content
def _call_vision(images, system, schema, prompt, api_key=None, model=None,
effort=None, max_tokens=None, labels=None):
"""Dispatch to the right provider's implementation.
Raises VisionError with a human-readable message on any failure callers
surface it in the UI rather than half-committing anything. Everything
provider-specific (request shape, auth, refusal/truncation handling,
usage-field names) lives in _call_anthropic / _call_openai below; this
only picks which one runs.
"""
if not images:
raise VisionError("No images to analyze.")
model = model if model in MODELS else DEFAULT_MODEL
caps = MODELS[model]
provider = caps.get("provider", "anthropic")
key = (api_key or _env_api_key(provider)) or None
if not key:
raise VisionError(
"No {} API key set. Add one in Settings, or export {} before "
"starting the app.".format(
"Anthropic" if provider == "anthropic" else "OpenAI",
"ANTHROPIC_API_KEY" if provider == "anthropic" else "OPENAI_API_KEY"))
if provider == "openai":
return _call_openai(images, system, schema, prompt, key, model,
max_tokens or MAX_TOKENS, labels)
return _call_anthropic(images, system, schema, prompt, key, model, caps,
effort or DEFAULT_EFFORT, max_tokens or MAX_TOKENS, labels)
def _call_anthropic(images, system, schema, prompt, key, model, caps, effort,
max_tokens, labels):
if anthropic is None:
raise VisionError(
"The anthropic package isn't installed. Run: "
"python3 -m pip install --user anthropic"
)
client = anthropic.Anthropic(api_key=key)
content = _build_content(images, labels, lambda mime, b64: {
"type": "image", "source": {"type": "base64", "media_type": mime, "data": b64}})
content.append({"type": "text", "text": prompt}) content.append({"type": "text", "text": prompt})
params = { params = {
"model": model, "model": model,
"max_tokens": max_tokens or MAX_TOKENS, "max_tokens": max_tokens,
"system": system, "system": system,
"output_config": {"format": {"type": "json_schema", "schema": schema}}, "output_config": {"format": {"type": "json_schema", "schema": schema}},
"messages": [{"role": "user", "content": content}], "messages": [{"role": "user", "content": content}],
@ -229,6 +281,69 @@ def _call_vision(images, system, schema, prompt, api_key=None, model=None,
} }
def _call_openai(images, system, schema, prompt, key, model, max_tokens, labels):
if openai is None:
raise VisionError(
"The openai package isn't installed. Run: "
"python3 -m pip install --user openai"
)
client = openai.OpenAI(api_key=key)
content = _build_content(images, labels, lambda mime, b64: {
"type": "image_url", "image_url": {"url": "data:{};base64,{}".format(mime, b64)}})
content.append({"type": "text", "text": prompt})
try:
response = client.chat.completions.create(
model=model,
messages=[{"role": "system", "content": system},
{"role": "user", "content": content}],
response_format={
"type": "json_schema",
"json_schema": {"name": "psa_grade_estimate", "strict": True, "schema": schema},
},
# Newer reasoning-capable models reject the older `max_tokens`
# name outright; this is the one Chat Completions accepts now.
max_completion_tokens=max_tokens,
)
except openai.AuthenticationError:
raise VisionError("OpenAI rejected the API key. Check it in Settings.")
except openai.PermissionDeniedError:
raise VisionError("That API key doesn't have access to {}.".format(model))
except openai.RateLimitError:
raise VisionError("OpenAI is rate-limiting you. Wait a moment and retry.")
except openai.BadRequestError as exc:
raise VisionError("OpenAI rejected the request: {}".format(exc))
except openai.APIConnectionError:
raise VisionError("Couldn't reach OpenAI. Check your connection.")
except openai.APIStatusError as exc:
raise VisionError("OpenAI error {}: {}".format(exc.status_code, exc))
choice = response.choices[0]
# content_filter is OpenAI's refusal equivalent; length is a truncation,
# same distinction Anthropic's stop_reason makes, different vocabulary.
if choice.finish_reason == "content_filter":
raise VisionError("OpenAI declined to analyze this image. Try a different photo.")
if choice.finish_reason == "length":
raise VisionError("The response was cut off. Try fewer/simpler images at a time.")
text = choice.message.content if choice.message else None
if not text:
raise VisionError("OpenAI returned no readable result for this image.")
try:
parsed = json.loads(text)
except ValueError:
raise VisionError("OpenAI's response wasn't valid JSON.")
usage = response.usage
return parsed, {
"input_tokens": getattr(usage, "prompt_tokens", None),
"output_tokens": getattr(usage, "completion_tokens", None),
"model": response.model,
}
GRADING_SYSTEM = """You estimate what PSA grade a trading card would likely \ GRADING_SYSTEM = """You estimate what PSA grade a trading card would likely \
receive, from photograph(s) of it. The card may be Pokemon, another trading \ receive, from photograph(s) of it. The card may be Pokemon, another trading \
card game, or a sports card PSA grades all of them on the same four \ card game, or a sports card PSA grades all of them on the same four \
@ -865,6 +980,7 @@ def price_guide():
rate_in, rate_out = current_rates(caps) rate_in, rate_out = current_rates(caps)
guide[model_id] = { guide[model_id] = {
"label": caps["label"], "label": caps["label"],
"provider": caps.get("provider", "anthropic"),
"supports_effort": caps["effort"], "supports_effort": caps["effort"],
"per_grade": round(est_in * rate_in / 1_000_000 "per_grade": round(est_in * rate_in / 1_000_000
+ est_out * rate_out / 1_000_000, 4), + est_out * rate_out / 1_000_000, 4),