For AI agents¶
This page is for AI agents (and the people building them) that call Sight on a
user's behalf: a coding assistant wiring Sight into an app, or an agent that
sends photos and reads results itself. The same content, in plain text for
tools, is at /llms.txt.
What Sight does, in one paragraph¶
Sight checks whether a piece of work was done right, from photos of it. Send
each photo with your own id for the job and for the photo; Sight checks it
against the account's template and answers with a verdict (pass,
fail, needs_review, incomplete), each check's outcome, the evidence
(which photo, which box, what was read), and an ask for anything that did
not pass: a sentence to show the user. Results can also be pushed to the app by
a signed webhook.
Rules an agent must follow¶
- The API key stays on a server. Never put it in a prompt, a log, a phone app, a web page or a repository. Read it from the environment.
- Always pass your own ids (
jobId,photoId). They make every call safe to repeat: a retry or a resend adds a photo once. - Send a link, not the photo. Pass a presigned https link to the photo in
the app's own storage (
imageUrl); do not download and forward it. pendingmeans wait, not "needs a person". UsewaitForJob,watchJobor the webhook; do not ask a user to review apendingcheck.- Show
askandstill_missingas they are. They are written for the user. Do not invent your own reasons for an outcome. - Verify every webhook before acting on it (
verifyWebhookon the raw body). Ignore repeats byevent.id; keep the delivery with the latestchecked_at. - Do not set or change the webhook unless the account's owner asked you to. It needs a key allowed to manage it, and every change is recorded.
- Ask
classes()before usingdetect. Only those names are known to the account's model. - On errors:
400means fix the request (readdetail);401/403means the key, not the request, is wrong: stop and tell the user;404means no such job or photo for this account;503is brief: retry with growing waits (the SDK already does).
Context to give an agent¶
Paste this into a system prompt or a tool description:
Sight (photo checks) API, https://sightapi.raku.so, key in X-Api-Key (server only).
SDK: npm @raku-technologies/sight -> new ChecksClient({ baseUrl, apiKey }).
- addPhoto({ jobId, template, slot, photoId, imageUrl, site?, readings?, detect?, wait? })
wait: omitted = answered at once (202); "fast" = rules + model now, readings "pending"; true = everything.
- getJob(jobId) | waitForJob(jobId) | watchJob(jobId) -> { verdict, checks[], still_missing[], checking, found[] }
- complete(jobId): no more photos; missing -> fail. removePhoto(jobId, photoId).
- listFailures(siteKey), classes(), annotatedPhoto(jobId, photoId) -> JPEG bytes.
- Webhook: setWebhook(url) (needs a permitted key; secret shown once), verifyWebhook(rawBody, header, secret).
verdict: pass | fail | needs_review | incomplete.
outcome: pass | fail | needs_review | not_assessable | missing | pending (pending = wait).
Always pass your own jobId and photoId; show `ask` to the user; never expose the key.
A tool definition¶
For an agent framework that takes JSON-schema tools, one tool covers sending a
photo; your code calls addPhoto with the arguments:
{
"name": "send_photo_for_checks",
"description": "Send one photo of a job to Sight to be checked. Returns the job's verdict, each check's outcome, and what is still missing.",
"input_schema": {
"type": "object",
"properties": {
"jobId": {"type": "string", "description": "The app's own id for the job"},
"template": {"type": "string", "description": "The check template's key"},
"slot": {"type": "string", "description": "Which photo this is, as the template names it"},
"photoId": {"type": "string", "description": "The app's own id for the photo"},
"imageUrl": {"type": "string", "description": "A presigned https link to the photo"},
"detect": {"type": "array", "items": {"type": "string"}, "description": "Optional: just these things to find"}
},
"required": ["jobId", "template", "photoId", "imageUrl"]
}
}
Where to read more¶
- Photo checks: the words and outcomes.
- Sending photos and Results back to your app.
- HTTP API: every route and field.
- The SDK's own README on npm.