API, REST contracts¶
Base URL: /api/v1
All responses are JSON.
Auth. Endpoints marked ๐ require a logged-in session: the vidit_session cookie (set by POST /auth/login, HttpOnly; Secure; SameSite=Lax) plus, for state-changing requests (POST/PUT/PATCH/DELETE), the X-CSRF-Token header echoing the JS-readable vidit_csrf cookie. There is no Authorization: Bearer flow; the cookie + CSRF pair is the only authenticated channel into the backend. Endpoints marked ๐ก๏ธ additionally require is_admin=true on the caller (returns 403 otherwise).
Transport security. Every response carries Strict-Transport-Security: max-age=15768000. The header has no includeSubDomains or preload directives.
Auth audit log. The /auth/* endpoints write to the auth_events table as a side-effect: login on success, failed_login on any rejected login (with user_id only when the address matched a live user), logout, register_pending (on POST /auth/register), register_resent (on POST /auth/resend-confirmation, on both the matched-pending and no-matching-pending branches so the rate-of-requests signal survives the always-204 discipline; user_id is always NULL since no user row exists yet), register_confirmed (on POST /auth/confirm-registration), password_reset_requested (on POST /auth/forgot-password, on both the known-email and unknown-email branches so the audit trail is a "rate of requests" signal), password_reset_completed, and password_changed (on POST /auth/change-password). Writes are best-effort inside a SAVEPOINT; an audit failure never breaks the auth flow.
Error envelope. Three shapes appear on the detail field of non-2xx responses, and frontend apiFetch (frontend/src/lib/api.ts) normalises all three. (1) Plain string, {"detail": "Invite code not found"} for direct HTTPException raises in routers (e.g. DELETE /admin/invite-codes/{id} 404). (2) Pydantic validation array, {"detail": [{"loc": [...], "msg": "...", "type": "..."}, ...]} for request-body / query-string validation failures (FastAPI default). (3) Typed envelope, {"detail": {"code": "<stable_id>", "message": "<human prose>"}} for business-rule errors raised from the service layer and translated by the router. Used by every /auth/register + /auth/confirm-registration + /auth/resend-confirmation error branch (codes: invalid_invite, email_already_registered, username_already_taken, email_pending_confirmation, username_pending_confirmation, invalid_or_expired_token), every /admin/* business-rule error branch (codes: user_not_found, geolocation_not_found, trust_reason_required, x_handle_conflict), and every POST /events, POST /events/requests, and POST /events/{id}/geolocate business-rule branch (codes: invalid_coordinates, too_many_files, media_required, invalid_proof, proof_image_required, tag_requirements_not_met, invalid_file, evidence_processing_failed, proof_files_mismatch, source_media_conflict; the create, request, and geolocate paths share the file/media codes via services/evidence_intake). POST /events/{id}/geolocate and POST /events/{id}/close add invalid_state when the row is not requested / detected. The code is the stable contract surface: branch on it, not on message. Status codes follow the per-endpoint contracts below.
Endpoints at a glance¶
Auth column: ๐ anonymous, ๐ logged-in, ๐ก๏ธ admin-only.
| Method | Path | Auth | Purpose |
|---|---|---|---|
| Auth | |||
| POST | /auth/register |
๐ | Stage a pending registration; sends confirmation email |
| POST | /auth/confirm-registration |
๐ | Confirm a pending registration (creates user, signs in) |
| GET | /auth/invites/{code}/check |
๐ | Advisory invite-code probe for the registration form |
| POST | /auth/resend-confirmation |
๐ | Re-send the confirmation email; invalidates previous token |
| POST | /auth/login |
๐ | Email + password โ session + CSRF cookies |
| POST | /auth/logout |
๐ | Clear session cookies (idempotent) |
| GET | /auth/me |
๐ | Current user |
| POST | /auth/forgot-password |
๐ | Email a single-use reset token (always 204) |
| POST | /auth/reset-password |
๐ | Consume reset token, set new password |
| POST | /auth/change-password |
๐ | Authenticated password rotation; requires current password |
| Events | |||
| GET | /events |
๐ | List one lifecycle view, located (default) or requested (ex /requests) |
| GET | /events/points |
๐ | Compact map-points tuples (cached) |
| GET | /events/possible-duplicates |
๐ | Soft-warning probe for the submit form |
| POST | /events/import-from-tweet |
๐ | Parse a tweet URL into a submit-form pre-fill payload |
| GET | /events/import-from-tweet/media |
๐ | Proxy fetch an X CDN media URL |
| POST | /events/import-archive/presign |
๐ | Mint a presigned direct-to-storage upload for your X data archive |
| POST | /events/import-archive |
๐ | Enqueue your staged archive (by upload_key) for the backfill worker |
| GET | /events/import-archive/{job_id} |
๐ | Poll your import job (status + assemble counts) |
| GET | /events/{id} |
๐ | Full event detail, any lifecycle state |
| POST | /events |
๐ | Create an event born geolocated (multipart, uploads media) |
| POST | /events/requests |
๐ | Open a request (multipart); creates a requested event (ex POST /requests) |
| DELETE | /events/{id} |
๐ | Owner-only hard delete + S3 sweep |
| POST | /events/{id}/geolocate |
๐ | Give an event a vouched location: requested | detected โ geolocated |
| POST | /events/{id}/close |
๐ | Owner withdraws a request or rejects a detection (โ closed) |
| POST | /events/{id}/investigate |
๐ | "I'm working on this" (idempotent, multi-analyst) |
| DELETE | /events/{id}/investigate |
๐ | Leave the working set |
| GET | /events/detections |
๐ | Your detected events awaiting a geolocate (paginated) |
| Search | |||
| GET | /search |
๐ | Free-text search across geolocations / requests / users |
| GET | /search/authors |
๐ | Username typeahead for the author filter |
| Tags | |||
| GET | /tags |
๐ | List tags (defaults to ones referenced by live geos) |
| POST | /tags |
๐ | Create a free tag (curated categories rejected) |
| Conflicts | |||
| GET | /conflicts |
๐ | List the conflict referential (?used=true narrows to conflicts on live events) |
| Users | |||
| GET | /users/{username} |
๐ | Public analyst profile |
| GET | /users/{username}/stats |
๐ | Aggregated shape of an analyst's work (status split, tags, activity) |
| PATCH | /users/me |
๐ | Edit your bio, avatar, external links |
| GET | /users/{username}/events |
๐ | List an analyst's geolocations |
| POST | /users/{username}/follow |
๐ | Follow (idempotent; self-follow โ 400) |
| DELETE | /users/{username}/follow |
๐ | Unfollow (idempotent; unknown user โ 404) |
| Timeline | |||
| GET | /timeline |
๐ | Activity feed from followed analysts |
| Webhooks | |||
| GET | /webhooks/x |
๐ | X webhook CRC challenge (HMAC answer, no DB) |
| POST | /webhooks/x |
๐ | X Account Activity delivery; signature-verified, queues bot mentions |
| Admin (collapsed below) | |||
| GET | /admin/me |
๐ก๏ธ | is_admin probe |
| GET | /admin/detection-stats |
๐ก๏ธ | Machine-extraction quality: reject-rate + pending missing-piece counts |
| POST/GET/DELETE | /admin/invite-codes[/{id}] |
๐ก๏ธ | Mint / list / revoke invite codes |
| GET | /admin/users |
๐ก๏ธ | Substring search on username/email |
| DELETE | /admin/users/{id} |
๐ก๏ธ | Soft delete (default) or ?hard=true GDPR erasure |
| DELETE | /admin/users/{id}/detected-events |
๐ก๏ธ | Purge every detected draft the user owns, account untouched |
| DELETE | /admin/events/{id} |
๐ก๏ธ | Soft delete or ?hard=true GDPR erasure |
| PATCH | /admin/users/{id}/trust |
๐ก๏ธ | Grant / revoke is_trusted + trust_reason |
| PATCH | /admin/users/{id}/x-handle |
๐ก๏ธ | Link / clear the bot-attribution X handle |
| POST/DELETE | /admin/seed-demo[-requests] |
๐ก๏ธ | Generate / drop demo geos + users / requests |
| POST | /admin/maintenance/reap-* |
๐ก๏ธ | Cron-style reapers (auth tokens, pending regs) |
Rate limits¶
One shared slowapi limiter (app/ratelimit.py), keyed per client IP, the right-most X-Forwarded-For entry (see engineering.md โ Particularities). Limits are per-endpoint; there is no global floor, so any endpoint absent from this table is unlimited. Buckets are in-process (one replica today). An over-quota request gets 429 with {"detail": "Rate limit exceeded. Try again later."}. RATE_LIMIT_ENABLED=false disables every limit at once (local dev). Every read limit in this table is behaviorally pinned (N requests answer, N+1 returns 429; see test_rate_limits.py); write limits have wiring-level coverage only.
| Endpoint | Limit (per IP) |
|---|---|
| Auth | |
POST /auth/login |
5/min + 30/hour |
POST /auth/register |
10/hour |
POST /auth/confirm-registration |
30/hour |
POST /auth/resend-confirmation |
5/hour |
POST /auth/forgot-password |
5/hour |
POST /auth/reset-password |
10/hour |
POST /auth/change-password |
10/hour (keyed per session) |
| Events | |
GET /events, GET /events/{id}, GET /events/detections |
120/min |
GET /events/points |
60/min |
GET /events/possible-duplicates |
60/min |
POST /events/import-from-tweet |
30/min |
GET /events/import-from-tweet/media |
60/min |
POST /events/import-archive/presign |
10/hour |
POST /events/import-archive |
10/hour |
GET /events/import-archive/{job_id} |
60/min |
POST /events, POST /events/requests, DELETE /events/{id} |
30/min |
POST /events/{id}/geolocate |
30/min |
POST /events/{id}/close, POST/DELETE /events/{id}/investigate |
60/min |
| Search / Tags | |
GET /search, GET /search/authors |
60/min |
GET /tags |
60/min |
POST /tags |
30/min |
GET /conflicts |
60/min |
| Users / Timeline | |
GET /users/{username}, GET /users/{username}/stats, GET /users/{username}/events, GET /timeline |
120/min |
PATCH /users/me |
30/min |
POST/DELETE /users/{username}/follow |
60/min |
| Admin ๐ก๏ธ | |
POST /admin/invite-codes ยท DELETE /admin/users/{id} ยท DELETE /admin/users/{id}/detected-events |
30/hour |
DELETE /admin/invite-codes/{id} ยท PATCH /admin/users/{id}/trust ยท PATCH /admin/users/{id}/x-handle ยท DELETE /admin/events/{id} |
60/hour |
POST/DELETE /admin/seed-demo[-requests] |
10/hour |
POST /admin/maintenance/reap-* |
30/hour |
GET /auth/invites/{code}/check and the read-only admin probes (GET /admin/me, /admin/detection-stats, /admin/users, /admin/invite-codes list) carry no limit. The /webhooks/x pair carries none either: the POST verifies the HMAC signature over the raw body (one HMAC, cheaper than any limiter bookkeeping), and the GET only ever signs tokens matching X's URL-safe CRC shape, the charset gate that keeps the responder from being a signing oracle for forged webhook bodies.
Auth¶
POST /auth/register¶
Stage a registration. Anonymous. No users row is created here: the submission lives in pending_registrations until the user proves they own the email by clicking the link in the confirmation message. The invite code is referenced by the pending row but not consumed; an abandoned signup does not burn the invite.
Request body:
{
"username": "kalush",
"email": "kalush@example.com",
"password": "โขโขโขโขโขโขโขโข",
"invite_code": "abc123"
}
Response 202:
No session cookie is set. The confirmation email is sent on a background task so the success and error branches return at the same wire timing.
Errors: | Code | Case | |------|------| | 400 | Invite code invalid, expired, revoked, or exhausted | | 409 | Email or username already registered (live or soft-deleted user) | | 409 | Email or username already has a live pending confirmation (distinct message) | | 429 | Rate-limited (10/hour/IP) |
POST /auth/confirm-registration¶
Anonymous. Consumes the token emailed by POST /auth/register, creates the users row, marks the invite consumed, and signs the analyst in (sets vidit_session + vidit_csrf cookies in the same response).
Request body:
Response 200: UserRead (same shape as GET /auth/me).
| Status | Meaning |
|---|---|
| 200 | Account created; cookies set; redirect to / |
| 400 | Token unknown, expired, or already consumed |
| 409 | Email or username was taken in the gap between register and confirm |
Rate-limited to 30/hour per IP.
GET /auth/invites/{code}/check¶
Anonymous. Pre-flight invite-code probe for the registration form. Mirrors the same validate_invite_code check the POST /auth/register step runs, so a 200 {"valid": true} here does not reserve the code, a concurrent registration can still consume it between the check and the submit.
Response 200:
Errors: | Code | Case | |------|------| | 404 | Invalid, exhausted, or expired invite code |
POST /auth/resend-confirmation¶
Anonymous. Re-mints the token for an outstanding pending registration and re-sends the confirmation email. Always 204 to avoid leaking which addresses are in flight. The previous token is invalidated by the re-mint, a shoulder-surfed link from the first email cannot be redeemed after the resend.
Request body:
Response 204 (always).
Rate-limited to 5/hour per IP.
POST /auth/login¶
Request body:
Response 200: UserRead (same shape as GET /auth/me). Sets the vidit_session HttpOnly cookie and the JS-readable vidit_csrf cookie.
Errors: | Code | Case | |------|------| | 401 | Wrong email or password | | 429 | Rate-limited (5/min/IP, 30/hour/IP) |
POST /auth/logout¶
Clears the session and CSRF cookies. Not session-gated (idempotent); like any mutating request it still requires the X-CSRF-Token header when a vidit_csrf cookie is present. Response 204: no body.
GET /auth/me ๐¶
Returns the current user.
Response 200:
{
"id": "uuid",
"username": "kalush",
"email": "kalush@example.com",
"is_trusted": false,
"trust_reason": null,
"bio": null,
"avatar_url": null,
"external_links": {},
"created_at": "2026-03-28T10:00:00Z"
}
is_trusted / trust_reason are public on purpose: the trust mark is a credibility signal, and the reason is what makes it credible. The profile fields (bio, avatar_url, external_links) ship with the self-payload so the sidebar avatar and "edit profile" form can render without a second fetch. is_admin is not on this shape; the admin role only surfaces via GET /admin/me. email_verified_at is not exposed: the pre-creation flow means there's no unverified-user state.
POST /auth/forgot-password¶
Anonymous. Emails a single-use reset token if the address matches an account. Always 204 to avoid user enumeration. Email-send failures are logged and swallowed for the same reason.
Body:
Response 204 (always, on success or unknown email).
Rate-limited to 5/hour per IP.
POST /auth/reset-password¶
Anonymous. Consumes a reset token and sets a new password. Tokens are single-use, expire PASSWORD_RESET_TOKEN_MINUTES after mint (default 15, the reset email quotes the same value), and become invalid the moment a fresh forgot-password is issued for the same user.
Body:
Response 204 on success.
| Status | Meaning |
|---|---|
| 204 | Password updated; client should redirect to /login |
| 400 | Token unknown, expired, already consumed, or wrong purpose, same opaque error to avoid leaking which |
Rate-limited to 10/hour per IP.
POST /auth/change-password ๐¶
Authenticated password rotation from the settings page. Requires re-asserting the current password so a stolen cookie can't lock the owner out. Audited as password_changed on success. After commit, a best-effort heads-up email goes to the address (no IP/UA, links to /forgot-password for owners who didn't trigger it). Email-send failure is swallowed (logged with user_id, never the address); the rotation succeeds either way.
Body:
Response 204 on success.
| Status | Meaning |
|---|---|
| 204 | Password updated; session cookie stays valid |
| 400 | Current password incorrect |
| 401 | Not authenticated |
| 422 | new_password shorter than 8 characters |
Rate-limited to 10/hour per session.
Events¶
GET /events¶
List one lifecycle view, newest first. Returns a lightweight card shape (no full proof).
Query params:
| Param | Type | Description |
|-------|------|-------------|
| view | string | located (default, the catalog: geolocated + detected rows, plus a closed row whose before_closed_status was detected) or requested (the open-call queue, ex /requests: requested rows, plus a closed row whose before_closed_status was requested). Anything else โ 422. |
| status | string (repeatable) | Narrows within the view, e.g. ?view=requested&status=closed. Repeat the param to OR within the bucket (?status=geolocated&status=detected). Values outside requested / detected / geolocated / closed return 422; a value the view can't contain returns an empty list. |
| conflict | string (repeatable) | Filter by conflict name, matched against the conflicts referential (conflicts.name), not tags. Repeat the param to OR within the conflict bucket (?conflict=Russian invasion of Ukraine&conflict=Gaza war). Combining with other buckets ANDs across them. |
| capture_source | string (repeatable) | Filter by capture-source tag name (?capture_source=Satellite&capture_source=Drone). Same semantics as conflict: OR within the bucket, AND across buckets, and the matched tag must carry category == "capture_source". |
| tag | string (repeatable) | Filter by tag name (any category). Repeat the param to OR within the tag bucket (?tag=drone&tag=tank). Combining buckets ANDs across them, the event must satisfy each bucket independently. |
| bbox | string | south,west,north,east (four comma-separated floats). 422 on malformed input, latitudes in [-90, 90], longitudes in [-180, 180], south โค north, west โค east. |
| event_date_from / event_date_to | date (YYYY-MM-DD) | Inclusive event-date range. Malformed values return 422 (used to silently 500 from Postgres InvalidDatetimeFormat). |
| submitted_from / submitted_to | date (YYYY-MM-DD) | Inclusive submission-date range. Same 422-on-malformed shape as the event-date filters. |
| author | string | Exact, case-insensitive match on owner username ("this analyst's work"; pick real handles via GET /search/authors). Whitelisted to [A-Za-z0-9_-]{1,50}, any other character returns 422. |
| limit | int | Default 200, must be in [1, 200], 422 otherwise. |
Response 200:
[
{
"id": "uuid",
"title": "Strike on depot, Donetsk",
"event_coords": { "lat": 48.123, "lng": 37.456 },
"event_date": "2026-03-15",
"is_demo": false,
"status": "geolocated",
"before_closed_status": null,
"owner": {
"id": "uuid",
"username": "kalush",
"is_trusted": false,
"trust_reason": null
},
"media": {
"id": "uuid",
"role": "source",
"storage_url": "https://โฆ/uploads/.../photo.jpg",
"media_type": "image"
},
"tags": [
{ "name": "Drone", "category": "capture_source" },
{ "name": "airstrike", "category": "free" }
],
"conflicts": [
{ "id": "uuid", "name": "Russian invasion of Ukraine", "wikidata_id": "Q110999040", "start_year": 2022, "end_year": null, "ongoing": true, "tier": "major" }
],
"investigator_count": null,
"investigators_sample": null
}
]
status is one of requested / detected / geolocated / closed; event_coords is null on a coordinate-less requested row. media is the card thumbnail: the event's source attachment, else its first proof image (null when it has neither; a proof video is never picked). The pick lives in backend/app/services/thumbnails.py, the one home every card surface uses. conflicts is the event's rows from the conflict referential (ConflictRead shape). investigator_count / investigators_sample (up to 3, newest first) populate only on view=requested, null on view=located. The same card shape flows through the profile feed, the timeline, and search hits.
GET /events/points¶
Compact [id, lat, lng, event_date, added_date, detected, demo] tuples for client-side clustering, no joins, no pagination. event_date / added_date are ISO YYYY-MM-DD (the created_at calendar day); event_date is null when unknown (the column is optional), and the map's event-date scrubber skips null-dated points instead of hiding them. The map buckets the dates for its timeline scrubbers and filters client-side. detected is 1 for a machine-detected row, 0 for a geolocated one; demo is 1 for a demo row, so the map's filter panel offers its hide-demo toggle only when one is present (flags, not status strings). Located rows only, so requested events never appear here.
Results are cached in-memory for 60s per unique filter combination; the response
echoes X-Cache: HIT|MISS and Cache-Control: public, max-age=30. Rate-limited
to 60/min/IP.
Query params: conflict, capture_source, tag, event_date_from, event_date_to
submitted_from, submitted_to, author (see GET /events for semantics), plus
map-only filters media (repeatable, ?media=image&media=video, matches a geolocation
carrying any attachment of a listed type; values are constrained to image/video, else
422), trusted_only (author is_trusted), and hide_demo (exclude demo rows). The date
params are still accepted but the map now filters dates client-side from the payload.
Response 200:
[
["6c1fโฆuuid", 48.123, 37.456, "2024-03-11", "2024-03-12", 0, 0],
["a0b2โฆuuid", 50.450, 30.523, "2024-05-02", "2024-05-04", 1, 0]
]
GET /events/possible-duplicates ๐¶
Soft-warning probe for the submit form: geolocations that might describe the same event. Never blocks submission (advisory only).
Match rule: within ~500 m geodesic of the proposed (lat, lng) AND (same source-URL host or same event_date). Auth-required; rate-limited to 60/min/IP.
Inputs are tolerated gracefully:
- Partial / scheme-stripped source URLs (
t.me/channel/12345) parse via a best-effort host extractor (prependshttp://and re-parses). Hosts that don't match^[a-z0-9][a-z0-9-]*(\.[a-z0-9-]+)+$after normalisation (lowercase, leadingwww.stripped) disable the host leg rather than 422-ing. - Malformed
event_datevalues disable the date leg. - If neither leg ends up usable, the response is
[](not an error) so the frontend can call this eagerly while fields are still being typed.
Query params:
| Field | Type | Required | Description |
|---|---|---|---|
| lat | float | yes | Latitude (-90 to 90) of the prospective submission. |
| lng | float | yes | Longitude (-180 to 180). |
| source_url | string | no | Best-effort host extracted for the host-match leg. |
| event_date | string (YYYY-MM-DD) | no | Compared exactly for the date-match leg. |
Response 200: array of up to 10 candidates ordered by distance ascending.
[
{
"id": "uuid",
"title": "Strike on depot, Donetsk",
"event_coords": { "lat": 48.123, "lng": 37.456 },
"event_date": "2026-05-01",
"source_url": "https://t.me/somechannel/12345",
"distance_m": 55.4,
"owner": {
"id": "uuid",
"username": "kalush",
"is_trusted": true,
"trust_reason": "Bellingcat affiliate"
}
}
]
POST /events/import-from-tweet ๐¶
Parse a public tweet URL into a pre-fill payload for the submit form. Read-only, never creates a row; the analyst submits the form. Rate-limited to 30/min/IP.
Data source is X's public syndication endpoint (the same backend the embeddable <blockquote class="twitter-tweet"> widget uses). It's unauthenticated and undocumented; the route surfaces upstream failures as 502 with a fixed error string the frontend renders verbatim ("Couldn't read tweet, fill the form manually"). Responses are cached in-memory for 1h per tweet ID to bound repeat fetches.
Request body:
Accepts both x.com and twitter.com (with or without www.), tolerates query string + fragment, and reduces the path to /<handle>/status/<id>. Anything else (profile, list, search, non-X host) returns 400.
Response 200:
{
"source_url": "https://x.com/source_handle/status/1234567890123456789",
"original_tweet_url": "https://x.com/analyst_handle/status/1234567890123456790",
"posted_at": "2025-11-12T14:33:00.000Z",
"author_handle": "analyst_handle",
"tweet_text": "<full OP tweet text>",
"suggested_title": "<first non-empty line, trimmed to 120 chars>",
"parsed_coords": [
{ "lat": 48.012345, "lng": 37.802411 }
],
"media": [
{ "kind": "video", "remote_url": "https://video.twimg.com/...", "content_type": "video/mp4", "origin": "quote" },
{ "kind": "image", "remote_url": "https://pbs.twimg.com/...", "content_type": "image/jpeg", "origin": "op" }
],
"quoted_tweet": {
"source_url": "https://x.com/source_handle/status/1234567890123456789",
"author_handle": "source_handle",
"tweet_text": "<full quoted tweet text>"
},
"detected": [
{
"lat": 48.012345,
"lng": 37.802411,
"title": "<derived title>",
"proof_text": "<cleaned tweet text>",
"detected_from_url": "https://x.com/analyst_handle/status/1234567890123456790",
"event_date": "2025-11-12",
"media": [
{ "kind": "image", "remote_url": "https://pbs.twimg.com/...", "content_type": "image/jpeg", "origin": "op" }
]
}
]
}
detected is the machine path's view of the same tweet, the DetectedGeolocs the assemble pipeline would produce, surfaced for inspection with zero DB writes (no row, no media fetch). One entry per parsed coordinate; empty when none parse. It's distinct from the human pre-fill above (parsed_coords + media): parsed_coords is candidates for the analyst to pick, detected is what the machine would persist as a detected row if this tweet were tagged or backfilled.
Every field is best-effort. parsed_coords runs four coordinate extractors (decimal, decimal + hemisphere, DMS, Google-Maps URL) over the OP then the quoted tweet, capped at 3 candidates. suggested_title is the OP's first usable line (leading hashtags / URLs / list markers / bare coordinates stripped), truncated to 120 chars on a word boundary; empty when nothing usable remains. media[].remote_url is always pbs.twimg.com or video.twimg.com.
source_url and source_posted_at are both nullable: they fill only on an explicit signal, never as a guess. See ingestion.md for the full contract shared with the machine detection path. source_url resolution priority:
- Quoted tweet's URL: when the OP quote-retweets, the quoted tweet is the source.
quoted_tweetcarries its metadata so the frontend can render the credit in the proof body.source_posted_atis the quote's post date. - First X / Telegram / YouTube link in the OP's
entities.urls: catches the OSINT convention of typingSource: https://t.me/<channel>/<id>in the body.source_posted_atstaysnull, the link carries no date. A coordinate link (Google Maps) or any other host is not a footage source.
Without either signal, source_url and source_posted_at are both null and the form field starts empty. The OP's own URL is never a fallback.
original_tweet_url is always the OP's canonical URL, kept separately so the proof body can credit the analyst even when source_url points at the source.
media[].origin (op = own attachment, quote = quoted tweet) is informational; the frontend routes by media type:
kind: "video"โ primary (lands infiles[]on the submit form).kind: "image"โ proof (loaded into the Tiptap proof body inline; it uploads as one of the create/geolocate multipart'sproof_files[]at publish, seePOST /events).- No video in the response โ no primary media is loaded; the analyst attaches the source media manually.
The syndication endpoint doesn't expose reply-chain media, so a video the analyst posted as a self-reply on the same thread is invisible to this route.
Errors: | Code | Case | |------|------| | 400 | Not a tweet URL (wrong host, profile / list / search path, malformed) | | 404 | Tweet not accessible (deleted, protected, never existed) | | 502 | Syndication endpoint timeout / 5xx / schema drift, frontend renders the "fill the form manually" banner |
GET /events/import-from-tweet/media?u=<url> ๐¶
Thin proxy that fetches a single X CDN media URL and streams the bytes back.
Auth-required. The u host is whitelisted to pbs.twimg.com / video.twimg.com, any other host returns 400 (SSRF guard). Per-stream byte cap (~110 MB) matches the upload pipeline's video ceiling plus HTTP framing overhead.
Query params:
| Field | Type | Required | Description |
|---|---|---|---|
| u | string | yes | Absolute X CDN URL (pbs.twimg.com / video.twimg.com, https:// only). |
Response 200: the upstream bytes, with the upstream Content-Type preserved and Cache-Control: private, max-age=300.
Errors:
| Code | Case |
|------|------|
| 400 | u host not in the whitelist |
| 404 | Upstream returned 404 |
| 502 | Upstream transport error, 5xx, or response above the size cap |
POST /events/import-archive/presign ๐¶
Step one of the archive import: mint a staging key and a presigned direct-to-storage upload for the caller's (browser-stripped) zip. The archive never transits the API. The target is an S3 POST policy (or the dev upload endpoint against local storage, same shape): POST a multipart/form-data form to upload.url carrying every upload.fields entry ahead of the file part, no credentials. The policy pins the exact key, application/zip, and the size guard (2 GB), and expires after 15 minutes. No content validation here.
Request: empty body.
Response 200:
{
"upload_key": "archive-imports/<user-id>/<uuid>.zip",
"upload": {
"url": "https://<bucket>.s3.<region>.amazonaws.com/",
"fields": { "key": "โฆ", "Content-Type": "application/zip", "policy": "โฆ", "โฆ": "โฆ" }
}
}
Errors: 401 not authenticated.
POST /events/import-archive ๐¶
Step two: enqueue the staged archive for the backfill worker. The upload is the consent: every geolocation lands detected, attributed to the caller (no handle-ownership check in this version). The request verifies the staged object (the caller's own upload_key, present, under the size guard; a storage HEAD, the zip is never opened here) and returns a queued job (202): the worker service (see ingestion.md) runs the import off the request path and emails the caller the outcome. Poll the job (below) for the counts. A malformed zip therefore surfaces as a failed job + failure email, not a synchronous 4xx; the browser strip catches the common shapes before upload.
Tweets-only intake guard. Only the allowlisted entries are extracted (tweets.js, tweets_media/); everything else (DMs, email, account data, deleted-*) is never read. Extraction is hardened against zip-slip and zip-bombs; the per-media caps at assemble time are the product limits (see ingestion.md).
Idempotent on (detected_from_url, coordinate), so a re-upload is a free catch-up. A detection with no recoverable media persists media-incomplete (the owner adds media before submitting).
A tweet that references its footage only through a linked status (Source: x.com/.../status/...) has that footage chased via syndication; an unreachable status still lands the tweet, just source-less. A tweet whose footage is a Telegram post (Source: t.me/<channel>/<id>) has that post's public embed chased for its date and, when the embed serves it, its media; a sensitive post degrades to link + date.
Request: JSON. upload_key from the presign; post_estimate (optional, โฅ 1) is the browser strip's cosmetic volume hint for the queued display (the worker stamps the exact totals).
Response 202:
{
"id": "uuid",
"status": "queued",
"post_estimate": 1240,
"progress_done": 0,
"progress_total": null,
"created": 0, "skipped": 0, "recreated": 0, "failed": 0,
"error": null,
"created_at": "2026-07-17T12:00:00Z",
"started_at": null,
"finished_at": null
}
Errors:
| Code | Case |
|------|------|
| 400 | archive_upload_invalid (not a staging key the caller minted: wrong shape, or another user's) |
| 401 | Not authenticated |
| 404 | archive_upload_missing (nothing uploaded at upload_key) |
| 413 | archive_too_large (the staged object is over the size guard) |
GET /events/import-archive/{job_id} ๐¶
One archive-import job, owner-only (someone else's job id reads as 404, indistinguishable from unknown). The upload page polls this until status is terminal; the completion email is the durable signal for an analyst who left.
status walks queued โ running โ done | failed. post_estimate is a free zip-metadata volume hint stamped at enqueue (declared tweets.js size over a per-record average; a display hint, not a promise); once the worker's parse has the exact detection count it stamps progress_total and batches progress_done as rows land, the upload page's live "137 / 412". The counts are final once done: created is new detected rows; skipped a pair a live row already held; recreated a previously rejected pair re-detected; failed a detection that raised mid-persist (the rest still land). A failed job keeps whatever landed before the failure (re-uploading skips it and continues); error is a terse operator-facing reason. Rate-limited to 60/min/IP.
Response 200: the job payload above, counts and timestamps filled per status.
GET /events/{id}¶
Full detail for a single event, in any lifecycle state.
Response 200:
{
"id": "uuid",
"title": "Strike on depot, Donetsk",
"event_coords": { "lat": 48.123, "lng": 37.456 },
"capture_source_coords": null,
"source_url": "https://t.me/channel/12345",
"proof": { "type": "doc", "content": [] },
"event_date": "2026-03-15",
"event_time": "14:30:00",
"source_posted_at": "2026-03-14T18:05:00Z",
"created_at": "2026-03-16T09:42:00Z",
"updated_at": "2026-03-16T09:42:00Z",
"requested_at": null,
"detected_at": null,
"geolocated_at": "2026-03-16T09:42:00Z",
"closed_at": null,
"is_demo": false,
"status": "geolocated",
"close_reason": null,
"before_closed_status": null,
"detected_from_url": null,
"detected_post_at": null,
"owner": {
"id": "uuid",
"username": "kalush"
},
"requested_by": null,
"geolocators": [
{ "id": "uuid", "username": "kalush", "is_trusted": false, "trust_reason": null }
],
"investigator_count": 0,
"investigators": [],
"media": [
{
"id": "uuid",
"role": "source",
"storage_url": "https://d10w3bld05vsky.cloudfront.net/uploads/.../video.mp4",
"media_type": "video",
"sha256": "f7c3bcd13f00e8a4b2d4e9b3f1a2c5d6e7f8901234567890abcdef1234567890",
"original_filename": "IMG_2034.MOV"
}
],
"thumbnail": {
"id": "uuid",
"role": "source",
"storage_url": "https://d10w3bld05vsky.cloudfront.net/uploads/.../video.mp4",
"media_type": "video",
"sha256": "f7c3bcd13f00e8a4b2d4e9b3f1a2c5d6e7f8901234567890abcdef1234567890",
"original_filename": "IMG_2034.MOV"
},
"tags": [
{ "name": "Drone", "category": "capture_source" }
],
"conflicts": [
{ "id": "uuid", "name": "Russian invasion of Ukraine", "wikidata_id": "Q110999040", "start_year": 2022, "end_year": null, "ongoing": true, "tier": "major" }
]
}
event_coords is the subject point, null on a coordinate-less requested event; every geolocated row carries it. capture_source_coords is the optional camera position, null unless the submitter set it. source_url / source_posted_at are null on a detected row with no declared source (see ingestion.md); a requested or geolocated row always carries a source_url. requested_by is the analyst who opened the request, null on a directly-created event (no request preceded it). geolocators is the durable credit list (who vouched the location, oldest first; empty until the first geolocate); investigators is the full "working on this" list (newest first, event_investigators) and investigator_count its length. close_reason / before_closed_status are null while the event is open. media carries only the event's source attachment(s); a proof image never appears here, it lives inline in the proof document as a URL. thumbnail is the picked card thumbnail (the source attachment, else the first proof image, else null; same rule as GET /events), so previews built on this payload (the map pin hover) render it without re-deriving the pick.
Errors: | Code | Case | |------|------| | 404 | Event not found |
POST /events ๐¶
Create an event directly, born geolocated. To open a request without coordinates, use POST /events/requests; to give an existing requested / detected event a location, use POST /events/{id}/geolocate.
Request body (multipart/form-data):
| Field | Type | Required | Description |
|-------|------|----------|-------------|
| title | string | yes | Title, 1-255 chars. |
| lat | float | yes | Latitude (-90 to 90) of the subject: what the footage shows. |
| lng | float | yes | Longitude (-180 to 180) of the subject. |
| capture_source_lat | float | no | Latitude of the camera position (where the footage was shot from). Both-or-neither with capture_source_lng. |
| capture_source_lng | float | no | Longitude of the camera position. |
| source_url | string | yes | Original source URL, โค2000 chars. |
| event_date | string (YYYY-MM-DD) | no | When the depicted event happened. Omitted / empty โ stored NULL (the footage doesn't always establish the date; renders as Unknown). |
| event_time | string (HH:MM) | no | Optional time-of-day for the event (UTC). Omitted / empty โ stored NULL. |
| source_posted_at | string (YYYY-MM-DDTHH:MM) | yes | When the source posted the media, a full instant, read as UTC. Required on this path; the analyst supplies it, since an off-platform source doesn't always carry a machine-readable date. Distinct from event_date and the submission time. |
| proof | string (JSON) | no | Serialized Tiptap document. Its inline images reference not-yet-uploaded files as placeholder://<filename>, resolved against proof_files. |
| tag_ids | string (JSON array) | yes | ["uuid1", "uuid2"]. Must include at least one capture_source tag (see Required categories below). |
| conflict_ids | string (JSON array) | yes | ["uuid1"]. Ids from the conflict referential. At least one is required (see Required categories below). |
| file | File | yes | Exactly one source file (image or video): the footage. |
| proof_files | File[] | no | The proof body's inline images, matched to its placeholder:// srcs by filename. At least one is required (see Required categories). |
Response 201: same shape as GET /events/{id}, born "status": "geolocated" with requested_by: null and the caller in geolocators.
Required categories. Three legs of the evidence floor, checked before any upload so a rejection doesn't pay an S3 round-trip: (1) exactly one source file; (2) at least one image in the proof body (an already-uploaded URL or a placeholder:// resolved from proof_files); (3) conflict_ids must resolve to at least one conflict (error message "A conflict is required") and tag_ids to at least one tag of category capture_source, the curated, server-managed taxonomy (see Tags). Both domains ship an escape value (conflict โ "Other", capture_source โ "Unknown") so the requirement is always satisfiable; either miss rejects with tag_requirements_not_met.
Errors:
| Code | Case |
|------|------|
| 400 | Typed {code, message} branch: invalid_coordinates, media_required (no source file), invalid_proof (sanitiser rejection), proof_image_required (no proof image), tag_requirements_not_met (missing conflict or capture_source tag), invalid_file (disallowed MIME / size), evidence_processing_failed, or proof_files_mismatch (a placeholder:// src with no matching proof_files upload, or vice versa) |
| 409 | source_media_conflict, a concurrent request raced past the one-source-per-event index |
| 413 | Request body exceeds the platform body-size cap (max_video_size + max_proof_images_per_event ร max_image_size + 10 MB headroom). Pre-checked by the HTTP-layer middleware before any bytes touch the worker; 413 responses traverse CORS so cross-origin callers see a clean status instead of a CORS error. |
| 422 | Malformed input: event_date (not a YYYY-MM-DD date), event_time (not HH:MM), source_posted_at (not an ISO datetime), more than max_proof_images_per_event files in proof_files (too_many_files), title over 255 chars, source_url over 2000 chars. All match the same-shape rejection on GET /events filter params and _parse_bbox. |
DELETE /events/{id} ๐¶
Owner-only delete. Cascades media, tag links, and contributor rows. A hard delete: the row and every S3 object it referenced are gone, distinct from the admin soft-delete (DELETE /admin/events/{id}) and from POST /events/{id}/close, which leaves the row readable.
Response 204: no body.
Errors: | Code | Case | |------|------| | 403 | Caller is not the owner | | 404 | Event not found (incl. soft-deleted) |
GET /events/detections ๐¶
The owner "Detections" queue: the caller's machine-detected events awaiting a geolocate, newest first (created_at desc). Scoped to current_user, it ignores any URL username and never exposes another analyst's rows. Powers /profile/{username}/detections, where the owner reviews and geolocates each detection. Returns the full detail shape (media + tags), not the lightweight list card, so the queue shows the evidence and computes geolocate-readiness (source media + a conflict + a capture_source tag) client-side without a per-row fetch.
Query params:
| Param | Type | Description |
|-------|------|-------------|
| page | int | Page number (default 1) |
| per_page | int | Results per page (default 20, max 100) |
Response 200: each item is the same shape as GET /events/{id}.
{
"items": [ { "id": "uuid", "status": "detected", "media": [], "tags": [] } ],
"total": 12,
"page": 1,
"per_page": 20
}
A detection carries no location it was promoted from; requested_by is always null here (a detection is machine-born, not opened as a request).
Errors: | Code | Case | |------|------| | 401 | Not authenticated |
POST /events/requests ๐¶
Open a request: creates a requested event with no coordinates yet (ex POST /requests). One source file is required, since the platform treats a request as an "unfinished geolocation"; coordinates, the camera point, tags, and the event date are all optional (an approximate guess is allowed, both-or-neither on each coordinate pair). The caller is recorded as both owner and requested_by; requested_by survives the later geolocate.
Request body (multipart/form-data):
| Field | Type | Required | Description |
|-------|------|----------|-------------|
| title | string | yes | Title; empty / whitespace-only rejected. Max 255 chars. |
| source_url | string | yes | URL where the media was found. Max 2000 chars. |
| proof | string (JSON) | no | In-progress proof (Tiptap document); sanitised server-side and image-free (no proof_files on this path, inline images are dropped by the sanitiser) |
| lat | float | no | Latitude of an approximate guess. Both-or-neither with lng. |
| lng | float | no | Longitude of an approximate guess. |
| capture_source_lat | float | no | Latitude of the camera position, if known. Both-or-neither with capture_source_lng. |
| capture_source_lng | float | no | Longitude of the camera position. |
| event_date | string (YYYY-MM-DD) | no | When the depicted event happened. Often unknown for a request. |
| event_time | string (HH:MM) | no | Optional time-of-day for the event (UTC); requires event_date. |
| source_posted_at | string (YYYY-MM-DDTHH:MM) | yes | When the source posted the media, a full instant (UTC). |
| tag_ids | string (JSON array) | no | ["uuid1", "uuid2"]. Not required to open a request; the curated floor is enforced at geolocate. |
| conflict_ids | string (JSON array) | no | Ids from the conflict referential. Optional here, like tag_ids. |
| file | File | yes | Exactly one source file (image or video). |
Response 201: same shape as GET /events/{id}, with "status": "requested" and event_coords / capture_source_coords null unless a guess was supplied.
Errors:
| Code | Case |
|------|------|
| 400 | Plain-string validation (empty / whitespace-only title or source_url) or a typed {code, message} branch: invalid_coordinates (a half-typed guess pair), media_required (no file), invalid_proof, invalid_file, evidence_processing_failed |
| 413 | Request body exceeds the platform body-size cap, same middleware as POST /events |
| 422 | title over 255 chars / source_url over 2000 chars, malformed event_date / event_time / source_posted_at, missing required source_posted_at, or event_time without event_date |
POST /events/{id}/geolocate ๐¶
Give an event a vouched location: transitions requested | detected โ geolocated in one atomic request, writing the caller's whole edited form. This is the single fulfil / geolocate path. A detected row is immutable machine output, so this is the only write to it, and it stays owner-only; a requested event is answerable by anyone, and the fulfiller becomes its owner (requested_by keeps the original poster). Multipart, mirroring POST /events: the form posts the whole row state and the server applies the field updates, media removals, and new-media uploads, then freezes the row as geolocated, in one transaction under a row lock (a concurrent geolocate on the same row serializes and the loser gets 409). Allowed only while requested / detected; a geolocated row is frozen.
Request body (multipart/form-data):
| Field | Type | Description |
|-------|------|-------------|
| title | string | 1-255 chars |
| lat | float | Latitude (-90 to 90) of the subject |
| lng | float | Longitude (-180 to 180) of the subject |
| capture_source_lat | float | Latitude of the camera position. Both-or-neither with capture_source_lng. |
| capture_source_lng | float | Longitude of the camera position. |
| source_url | string | โค2000 chars, the footage origin. A detected draft may start with no declared source (null, see ingestion.md): a blank value here 400s as source_url_required, since a geolocated row always carries one. Fulfilling a requested event ignores this field and keeps the request's source_url (a fulfiller must not rewrite the requester's evidence anchor) |
| event_date | string (YYYY-MM-DD) | When the depicted event happened. Optional, mirroring create: empty / omitted stores NULL (renders as Unknown) |
| event_time | string (HH:MM) | Optional time-of-day for the event (UTC); empty / omitted clears it |
| source_posted_at | string (YYYY-MM-DDTHH:MM) | When the source posted the media, a full instant (UTC). Required on this path; the analyst supplies it, since an off-platform source doesn't always carry a machine-readable date |
| proof | JSON string | Tiptap document (sanitised); its placeholder:// srcs resolve against proof_files, already-uploaded URLs pass through untouched |
| tag_ids | JSON string (UUID[]) | Replaces the tag set wholesale |
| conflict_ids | JSON string (UUID[]) | Replaces the event's conflict set wholesale |
| remove_media_ids | JSON string (UUID[]) | Existing source media to drop (S3 swept) |
| files | file[] | New source media to add (0 or 1; kept + new must total exactly one, same allowlist + size limits as create) |
| proof_files | file[] | New proof images referenced by placeholder:// srcs in proof |
detected_from_url (the provenance anchor, the post the detection was imported from) and status carry no field, so a caller that sends them is ignored. Blocked until the evidence floor a direct create meets is satisfied by the post-geolocate state: exactly one source media (kept + new), at least one proof image in the final proof body, and one conflict + one capture_source tag. A requested event and a machine detection are both born without the curated floor, so it is enforced here; the fulfiller adds the conflict and tags as part of the geolocate.
Response 200: same shape as GET /events/{id} (now "status": "geolocated", the caller added to geolocators).
Errors:
| Code | Case |
|------|------|
| 400 | invalid_coordinates, invalid_proof, proof_image_required (no proof image in the final body), tag_requirements_not_met, a rejected file (invalid_file / evidence_processing_failed), no surviving source media (media_required), proof_files_mismatch, or source_url_required (a detected draft with no declared source, geolocated with a blank source_url field) |
| 403 | Caller is not the owner of a detected draft (a requested event is answerable by anyone) |
| 404 | Event not found (incl. soft-deleted) |
| 409 | Row is not requested / detected (invalid_state, a geolocated row is frozen), or source_media_conflict (a concurrent edit raced past the one-source cap) |
| 422 | Kept + new source media over one (too_many_files), or more than max_proof_images_per_event proof files |
POST /events/{id}/close ๐¶
Close an event: withdraw a requested row or reject a detected draft, owner-only, in one verb. The row stays publicly visible (transparency: a queue entry that didn't produce a geolocation, or a machine draft judged wrong); before_closed_status records which state it left (drives the status badge and the requested-view routing). A closed detected row is re-importable if the same tweet is later re-detected. Distinct from DELETE, which removes the row for good.
Request body:
close_reason is required (1-2000 chars) and stays publicly visible on the closed row.
Response 200: same shape as GET /events/{id} (now "status": "closed").
Errors:
| Code | Case |
|------|------|
| 403 | Caller is not the owner |
| 404 | Event not found (incl. soft-deleted) |
| 409 | Row is not requested / detected (invalid_state, geolocated and closed are both terminal here) |
| 422 | close_reason missing or over 2000 chars |
POST /events/{id}/investigate ๐¶
Signal "I'm working on this" on a requested event. Multi-analyst: several investigators can hold the signal on one event at once, it's a coordination hint, not a single-claimer reservation. Idempotent, re-signalling is a 204 no-op, not a 409.
Response 204: no body.
Errors:
| Code | Case |
|------|------|
| 404 | Event not found (incl. soft-deleted) |
| 409 | Event status is not requested |
DELETE /events/{id}/investigate ๐¶
Caller leaves the working set. Idempotent (204 even if the caller wasn't signalling).
Response 204: no body.
Errors:
| Code | Case |
|------|------|
| 404 | Event not found (incl. soft-deleted) |
| 409 | Event status is not requested (a terminated event's signals are frozen history) |
Requests, geolocations, and detections are /events views¶
There is no /requests router. A request is a requested event, a geolocation is a geolocated event, and a detection is a detected event, all rows on the one events table, distinguished only by status. Every read and write above already covers all three:
- List / detail:
GET /events(view=requestedis the request queue,view=locatedthe geolocation catalog, both carrydetectedrows too) andGET /events/{id}(any status). - Open a request:
POST /events/requests(no coordinates required). - Fulfil a request, or vouch a detection:
POST /events/{id}/geolocate(requested|detectedโgeolocated, one verb for both). - Withdraw a request, or reject a detection:
POST /events/{id}/close(one verb for both,before_closed_statustells them apart). - "I'm working on this" on a request:
POST/DELETE /events/{id}/investigate. - Remove:
DELETE /events/{id}(owner hard delete) orDELETE /admin/events/{id}(admin soft/hard delete).
GET /events?view=requested cards additionally carry investigator_count / investigators_sample; GET /events/{id} always carries the full investigators list and geolocators. Search groups a hit under requests when its status is requested, see below.
Search¶
Slice-1 full-text discovery surface across the three first-class entity types. Backed by two Postgres GIN indexes on to_tsvector('simple', โฆ) expressions: one over events.title and one over users.username || ' ' || users.bio (migration o1j3k5l7m9n1). One FTS query path over the single events table; the located (geolocations) and requested (requests) groups run the same title index with different WHEREs (status IN ('geolocated', 'detected') AND event_coords IS NOT NULL vs status = 'requested'). The simple dictionary keeps matching predictable. The response is still grouped by entity type.
Out of scope for slice 1: searching source_url, JSONB-content search (events.proof), per-group infinite scroll, and the filter chips beyond the entity-type pick.
GET /search ๐¶
Query params:
| Param | Type | Description |
|-------|------|-------------|
| q | string | Free-text query. Empty / whitespace-only short-circuits to empty groups (unless a filter is active). |
| type | enum | all (default), event (the two event groups: what the search page's unified "Events" chip sends), geolocation, request, or user. Anything else โ 422. |
| limit | int | Per-group cap. 1 โค limit โค 50, default 20. |
| filter set | | The standard event filter set, same names and semantics as GET /events: status, conflict, capture_source, tag, media (repeatable), event_date_from / event_date_to, submitted_from / submitted_to, author, trusted_only, hide_demo. Scopes the two event groups (a status value a group's view can't contain empties that group). |
Any active filter empties the users group (the filters are event predicates; an unfiltered analyst list next to a filtered event view would read as if the filter applied). With an empty q and at least one active filter, browse mode: the filtered view, newest first, plain titles as their own highlight (the profile's "Show more" entry point); typing then narrows within it.
Ranking: ts_rank descending then created_at descending as a stable tie-breaker.
Soft-delete: every group filters deleted_at IS NULL at query time.
Highlight markers: each hit carries one or more *_highlight fields with STX (U+0002) / ETX (U+0003) control bytes around matched fragments. JSON encodes them as /. The frontend (lib/search.ts::splitHighlights) splits on those bytes and wraps the inner segments in <mark>, no raw HTML crosses the wire, so it's XSS-safe by construction.
Response 200:
{
"geolocations": [
{
"id": "uuid",
"title": "Strike on warehouse complex, Donetsk Oblast",
"title_highlight": "Strike on warehouse complex, Donetsk Oblast",
"lat": 48.01, "lng": 37.80,
"event_date": "2026-04-15",
"is_demo": false,
"status": "geolocated",
"owner": { "id": "uuid", "username": "osint_analyst", "is_trusted": true, "trust_reason": "โฆ" },
"media": [{ "id": "uuid", "role": "source", "storage_url": "โฆ", "media_type": "image" }],
"tags": [{ "id": "uuid", "name": "airstrike", "category": "free" }]
}
],
"requests": [
{
"id": "uuid",
"title": "Footage from Kharkiv area, can someone place it?",
"title_highlight": "Footage from Kharkiv area, can someone place it?",
"source_url": "https://twitter.com/โฆ",
"status": "requested",
"created_at": "2026-04-12T08:00:00Z",
"is_demo": false,
"owner": { "โฆ": "โฆ" },
"media": [{ "id": "uuid", "storage_url": "โฆ", "media_type": "image" }],
"tags": [],
"claimer_count": 3
}
],
"users": [
{
"id": "uuid",
"username": "kharkiv_osint",
"username_highlight": "kharkiv_osint",
"bio": "Tracking armoured movement in Eastern Ukraine.",
"bio_highlight": null,
"is_trusted": true,
"trust_reason": "Established public track record",
"avatar_url": null
}
],
"total": { "geolocations": 1, "requests": 1, "users": 1 },
"query": "kharkiv",
"type": "all"
}
media on both event groups carries the picked card thumbnail (at most one row: the source attachment, else the first proof image), the same rule as the GET /events card.
bio_highlight is null when only the username matched, the UI uses this to hide the snippet block instead of rendering an un-highlighted bio. Groups the caller didn't request via type= come back as empty arrays.
total is a fixed-key object (geolocations, requests, users), each the pre-LIMIT match count for its group (so the UI renders "3 of 142", not "3 of 3"). type echoes the request and is one of all, geolocation, request, user.
Errors:
| Code | Case |
|------|------|
| 401 | Unauthenticated |
| 422 | type outside the allowed set, or limit outside [1, 50] |
GET /search/authors¶
Username typeahead for the author filter (the map's and the search page's Author section). The author filter is an exact match, so this picker is how a partial name becomes a real handle: case-insensitive substring over live users, prefix matches first then alphabetical, capped at 8. q takes the same [A-Za-z0-9_-]{1,50} gate as ?author= (empty returns an empty list; anything else 422). Rate-limited to 60/min/IP.
Response 200:
Tags¶
GET /tags¶
List tags. By default returns only tags referenced by at least one live geolocation.
Query params:
| Param | Type | Description |
|-------|------|-------------|
| category | string | capture_source or free |
| curated | bool | When true, return the full curated capture_source taxonomy regardless of live usage, ignoring the default usage filter. Conflicts are no longer tags; the full conflict list lives on GET /conflicts. |
Response 200:
[
{ "id": "uuid", "name": "Drone", "category": "capture_source" },
{ "id": "uuid", "name": "airstrike", "category": "free" }
]
POST /tags ๐¶
Create a tag. Only free tags are creatable; capture_source is server-managed and rejected with 403.
Request body:
Validation. name is stripped of leading / trailing whitespace before any check or DB write, then bounded 1 <= len(name) <= 100 (the String(100) column cap on tags.name). Empty or whitespace-only names return 422. Duplicate-name detection is case-sensitive to match the DB unique constraint: Drone and drone are distinct rows, so two analysts using different casing will create two tags.
Response 201:
Response 403: category is not free.
Response 409: a tag with the same name already exists.
Errors: | Code | Case | |------|------| | 409 | A tag with this name already exists |
Conflicts¶
GET /conflicts¶
List the conflict referential, ordered ongoing first then by name. Server-managed (the daily Wikipedia sync, the one-shot Wikidata seed, operator rows; see ingestion.md): there is no create endpoint. The default returns every row, ongoing and ended alike, so the submit picker can offer ended conflicts for archival footage. Rate-limited to 60/min/IP.
Query params:
| Param | Type | Description |
|-------|------|-------------|
| used | bool | When true, return only conflicts carried by at least one live event, so a filter UI never surfaces a chip that matches zero results. Mirrors the default orphan filtering on GET /tags. |
Response 200:
[
{ "id": "uuid", "name": "Russian invasion of Ukraine", "wikidata_id": "Q110999040", "start_year": 2022, "end_year": null, "ongoing": true, "tier": "major" },
{ "id": "uuid", "name": "Western Sahara conflict", "wikidata_id": "Q1152920", "start_year": 1970, "end_year": null, "ongoing": false, "tier": null }
]
start_year / end_year disambiguate same-named historical entries. tier is the Wikipedia death-toll tier (major, minor, conflict; see data-model.md), NULL for rows the sync has never classified; clients use it to rank the default picker list. last_seen_at and source are sync internals and stay off the wire.
Ongoing-conflict names and dates derive from Wikipedia's "List of ongoing armed conflicts", available under CC BY-SA 4.0; any surface listing them should carry that attribution.
Users¶
GET /users/{username}¶
Public profile of an analyst.
Response 200:
{
"id": "uuid",
"username": "kalush",
"is_trusted": false,
"trust_reason": null,
"bio": "OSINT analyst tracking armoured movement in Eastern Ukraine.",
"avatar_url": "https://cdn.example.com/avatars/kalush.jpg",
"external_links": {
"x": "@kalush",
"discord": null,
"website": "https://kalush.example.com",
"github": null
},
"created_at": "2026-03-28T10:00:00Z",
"geolocations_count": 42,
"followers_count": 17,
"following_count": 5,
"is_following": false
}
is_trusted toggles via PATCH /admin/users/{id}/trust; trust_reason is required when granting. bio / avatar_url / external_links are self-set via PATCH /users/me, defaults are null / null / {}. is_following is true only when the caller is authenticated and follows this user; anonymous viewers and self-views always get false. Email is never on this shape.
Errors: | Code | Case | |------|------| | 404 | User not found |
GET /users/{username}/stats¶
Aggregated shape of an analyst's work, over their live events only (deleted_at IS NULL). Pure aggregation over existing columns; drives the profile's insights section (status split, media volume, top theatres, capture lens, 12-month activity).
Response 200:
{
"geolocated_count": 2,
"detected_count": 1,
"closed_count": 1,
"total_events": 4,
"media_count": 2,
"top_conflicts": [{ "name": "Russo-Ukrainian War", "count": 2 }],
"capture_sources": [{ "name": "dashcam", "count": 1 }],
"monthly_activity": [{ "month": "2025-08", "count": 0 }, { "month": "2025-09", "count": 3 }]
}
total_events is the sum of the three status counts (requested events are not part of the split). top_conflicts and capture_sources are capped at 5, ordered by count desc then name. monthly_activity buckets event_date into the last 12 calendar months (current month last), zero-filled.
Errors: | Code | Case | |------|------| | 404 | User not found |
PATCH /users/me ๐¶
Edit your own profile, bio, avatar URL, and Linktree-style external account handles.
Body (all fields optional; absent = leave column alone, explicit null or empty string = clear):
{
"bio": "OSINT analyst, Eastern Ukraine armoured movement.",
"avatar_url": "https://cdn.example.com/avatars/me.jpg",
"external_links": {
"x": "@me",
"discord": "me",
"website": "https://me.example.com",
"github": "@me"
}
}
bio is capped at 500 characters. avatar_url must be http or https, javascript: and other schemes are rejected (XSS class blocked at write time). external_links is wholesale-replaced, not deep-merged: send the full panel each time. Per-platform values are 200 chars max; values are free-form strings (handle or URL).
Response 200: the updated UserRead (same shape as GET /auth/me).
Errors: | Code | Case | |------|------| | 401 | Not authenticated | | 422 | Validation failure (bio too long, non-http(s) URL, unknown field) |
GET /users/{username}/events¶
Geolocations for a given analyst.
Query params:
| Param | Type | Description |
|-------|------|-------------|
| page | int | Page number (default 1) |
| per_page | int | Results per page (default 20, max 100) |
Response 200:
{
"items": [
{
"id": "uuid",
"title": "Strike on depot, Donetsk",
"lat": 48.123,
"lng": 37.456,
"event_date": "2026-03-15",
"media": { "id": "uuid", "storage_url": "https://โฆ/abc.jpg", "media_type": "image" },
"tags": [{ "name": "Drone", "category": "capture_source" }],
"conflicts": [{ "id": "uuid", "name": "Russian invasion of Ukraine", "wikidata_id": "Q110999040", "start_year": 2022, "end_year": null, "ongoing": true, "tier": "major" }]
}
],
"total": 42,
"page": 1,
"per_page": 20
}
media is the picked card thumbnail (same rule as GET /events), null when the event has neither a source attachment nor a proof image; the full media list is on the detail payload only.
POST /users/{username}/follow ๐¶
Follow another analyst. Idempotent, re-following a user you already follow returns 204 without error. Self-follow is rejected with 400.
Response 204: no body.
Errors: | Code | Case | |------|------| | 400 | Cannot follow yourself | | 401 | Not authenticated | | 404 | Target user not found or soft-deleted |
DELETE /users/{username}/follow ๐¶
Unfollow another analyst. Idempotent. Unknown username returns 404 rather than no-op'ing.
Response 204: no body.
Errors: | Code | Case | |------|------| | 401 | Not authenticated | | 404 | Target user not found or soft-deleted |
Timeline¶
GET /timeline ๐¶
Activity feed of geolocations submitted by analysts the current user follows, ordered by event date descending.
Query params:
| Param | Type | Description |
|-------|------|-------------|
| page | int | Page number (default 1) |
| per_page | int | Results per page (default 20, max 100) |
Response 200: same PaginatedEvents shape as GET /users/{username}/events.
Errors: | Code | Case | |------|------| | 401 | Not authenticated |
Admin¶
All routes below are mounted under /admin and gated by the require_admin FastAPI dependency. require_admin layers on top of get_current_user, so a deactivated admin (is_active=false) loses access immediately.
16 admin endpoints, rarely-touched ops surface (invites, detection-quality metrics, soft/hard delete, trust toggle, X handle link, demo seeding, maintenance reapers). Expand for full contracts.
### `GET /admin/me` ๐ก๏ธ **Response 200:** Returns 403 for non-admins, 401 for anonymous callers. ### `GET /admin/detection-stats` ๐ก๏ธ Quality signal on the machine-extraction pipeline. A **machine detection** is an event imported from X (the archive backfill or the bot), identified by `detected_from_url` being set; a human submit always carries `detected_from_url = null`. Demo rows (`is_demo`) are excluded from both aggregates. Read-only, no audit row (a metric read is not an administrative act). **Reject-rate** is the share of machine detections dismissed while still a draft, whichever door they left through. A machine detection counts as a reject if either an owner closed it straight out of `detected` (`status = "closed"` with `before_closed_status = "detected"`) or an admin soft-deleted it while it was still `detected` (`deleted_at` set with `status = "detected"`). A detection the owner vouched (promoted to `geolocated`) is **not** a reject, even once soft-deleted (it was vouched before removal); one still awaiting review is **not** a reject yet. `reject_rate` is `machine_rejected / machine_total` as a 0..1 ratio (`0` when there are no machine detections). Counted over every (non-demo) machine row, soft-deleted or not: the metric measures what the pipeline produced. Two counting edges the metric accepts, both favouring over-counting dismissals over under-counting them: an owner **hard-delete** (`DELETE /events/{id}` on an own draft) removes the row from both counts entirely; an **account-departure cascade** soft-delete counts that account's pending drafts as rejects. The `pending_*` counts profile the **live** `detected` queue (`deleted_at IS NULL`, machine rows only, demo excluded): drafts missing a piece the geolocate floor will demand (a source media, a proof-role image, or a `source_url`), so a low-quality extraction run is visible before an analyst opens the queue. **Response 200:**{
"machine_total": 420,
"machine_rejected": 37,
"reject_rate": 0.088,
"pending": 61,
"pending_missing_source_media": 4,
"pending_missing_proof_image": 9,
"pending_missing_source_url": 12
}
{
"id": "8e67f0โฆ",
"code": "abc123xyz",
"max_uses": 1,
"use_count": 0,
"expires_at": "2026-05-23T10:00:00Z",
"revoked_at": null,
"created_at": "2026-05-09T10:00:00Z",
"status": "active",
"redeemer": null,
"used_at": null,
"x_handle": "osint_hawk"
}
[
{
"id": "โฆ", "code": "โฆ", "status": "exhausted", "max_uses": 1, "use_count": 1,
"expires_at": null, "revoked_at": null, "created_at": "โฆ", "used_at": "โฆ", "x_handle": "osint_hawk",
"redeemer": {
"user_id": "โฆ",
"username": "osint_hawk",
"email": "hawk@example.com",
"is_admin": false,
"is_trusted": false,
"trust_reason": null,
"x_handle": "osint_hawk",
"archives_imported": 1,
"bot_detection_count": 3,
"detected_count": 12,
"geolocated_count": 4,
"last_login_at": "โฆ"
}
}
]
[
{
"id": "โฆ",
"username": "tester2",
"email": "tester2@example.com",
"is_admin": false,
"is_trusted": true,
"trust_reason": "Established OSINT track record",
"x_handle": "tester2",
"created_at": "โฆ"
}
]
{
"user_id": "โฆ",
"username": "throwaway",
"mode": "soft",
"deleted_at": "2026-05-09T16:45:00Z",
"cascaded_geolocations": 5,
"media_count": 0
}
{
"geolocation_id": "โฆ",
"title": "Strike on depot, Donetsk",
"mode": "soft",
"deleted_at": "2026-05-09T16:30:00Z",
"media_count": 0
}
{
"created": 20,
"templates": 3,
"authors": 5,
"with_claims": 11,
"open": 12,
"fulfilled": 5,
"closed": 3
}
Webhooks¶
The X Account Activity webhook, the bot's nominal mention delivery (see ingestion.md). Unauthenticated by design: X calls it, and the HMAC signature over the raw body (the app's consumer secret, held only by X and the deployment) is the gate.
GET /webhooks/x¶
X's Challenge-Response Check (CRC), sent at registration and then hourly; a wrong or slow answer deactivates the webhook. Answered in-request, no DB.
Query: crc_token (required). Must match ^[A-Za-z0-9_-]{1,200}$ (X's CRC tokens are short URL-safe strings). The gate is what keeps the responder from being a signing oracle: the answer is the exact HMAC construction the POST verifies over the raw body, and a JSON webhook body can never fit that charset.
Response 200:
Response 400: crc_token outside the URL-safe shape.
Response 503: the X credentials are not configured on this deployment.
POST /webhooks/x¶
One Account Activity delivery. The x-twitter-webhooks-signature header must carry sha256=<base64(HMAC-SHA256(consumer_secret, raw_body))>; compared constant-time as bytes, mismatch โ 401. A body over 512 KiB โ 413 before the body is read (an AAA delivery is small). A valid signature always answers 200, whatever the payload: a foreign for_user_id, non-mention events, or the bot's own posts are ignored (a non-2xx would make X retry and eventually deactivate the webhook). Mentions are reduced to the internal shape and queued in bot_webhook_events; the import worker runs the pipeline, never the request.
Response 200:
Response 503: the consumer secret or the bot user id is not configured (an empty bot user id would otherwise silently drop every delivery).
General conventions¶
Pagination¶
Endpoints that return paginated lists use this shape:
GET /events and GET /events/points are intentionally unpaginated today. A hard server-side LIMIT will land before public read access.
Errors¶
All errors follow this shape:
File limits¶
| Type | Extensions | Max size |
|---|---|---|
| Image | jpg, png, webp | 10 MB |
| Video | mp4, webm | 95 MiB |