HR-ATS-Portal/Sync_read.md

229 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters!

This file contains ambiguous Unicode characters that may be confused with others in your current locale. If your use case is intentional and legitimate, you can safely ignore this warning. Use the Escape button to highlight these characters.

# Read-status sync (`/sync/*`)
Tracks **which messages got read or unread** — and which were deleted — without
re-downloading the mailbox. It sits on Microsoft Graph's **delta query**: Graph
hands you a cursor, and every later call with that cursor returns *only* what
changed since it was issued.
Five of the six endpoints share one piece of state: a delta cursor per
**(signed-in user + folder)**, persisted to disk so a restart doesn't re-backfill
the whole folder. The sixth — the per-message lookup — is deliberately outside
that machinery: it reads one id live and touches no cursor.
| Method | Path | Purpose |
| ------ | ---- | ------- |
| GET | `/sync/read-status` | Run one sync round **now** (synchronous) |
| GET | `/sync/read-status/changes` | Replay the last round's **full** result |
| GET | `/sync/read-status/message/{id}` | One message's status, by id — cursor-free |
| GET | `/sync/read-status/status` | Watcher health + cursor state |
| POST | `/sync/read-status/watch` | Start the background poller |
| DELETE | `/sync/read-status/watch` | Stop the background poller |
All require the API bearer token, and act on the **signed-in user's** mailbox —
they answer `401` until device-code sign-in completes.
---
## `GET /sync/read-status`
The workhorse. Asks Graph "what changed in this folder since my cursor?", emits
the changes, and advances the cursor.
The **first** call has no cursor, so it backfills the entire folder — an Inbox
with 4,700 messages is 47 pages of 100. Every call after that is incremental and
usually near-empty.
| Param | Default | Meaning |
| ----- | ------- | ------- |
| `folder` | `inbox` | Well-known name (`inbox`, `sentitems`, …) or folder id. Graph delta is **folder-scoped** — there is no all-mail delta |
| `since` | | ISO8601 lower bound, **initial sync only** (`receivedDateTime ge …`). The way to keep a first backfill small |
| `reset` | `false` | Discard the saved cursor and start a fresh baseline |
| `max_pages` | `10` | Cap on Graph pages (100 msgs each) fetched **per call** |
| `limit` | `10` | Cap on messages returned **in this response** |
`max_pages` and `limit` are independent and easy to confuse:
- **`max_pages` bounds the work.** Hit the cap and the call returns
`complete: false`, having saved its position; the next call resumes exactly
where it stopped. No changes are skipped, and no cursor is written until the
backfill genuinely finishes.
- **`limit` only trims the JSON.** It has no effect on how much is fetched.
`count` stays the true total, and the untruncated set is on
`/sync/read-status/changes`.
```jsonc
{
"synced_at": "2026-08-07T10:15:00Z",
"folder": "inbox",
"count": 1000, // changed messages this call actually fetched
"removed_count": 0, // deleted / moved out of the folder
"initial_sync": true, // this round is part of the first backfill
"complete": false, // hit max_pages — call again to continue
"pages": 10, // Graph pages fetched by this call
"truncated": true, // limit cut the lists below
"value": [ { "id": "AAMk…", "isRead": true,
"lastModifiedDateTime": "2026-08-07T10:14:52Z",
"subject": "Invoice #421" } ],
"removed": [ { "id": "AAMk…", "reason": "deleted" } ]
}
```
`value` is sorted newest-modified first before `limit` is applied, so a
truncated response shows the most recent changes rather than an arbitrary slice.
Only the four `$select` fields above come back — this endpoint is about *status*,
not content; use `GET /emails/{id}` for bodies.
## `GET /sync/read-status/changes`
Read-only replay of whatever the **last** round produced. No Graph call, cursor
untouched, safe to hit repeatedly.
Two reasons it exists:
1. It holds the **untruncated** lists — this is how you get the other 990 items
when `limit` trimmed the response.
2. It's the only way to collect what the **background watcher** found, since the
watcher has no caller to return to.
`404` until some sync has run. One buffer, last-writer-wins: the next round
overwrites it, so with the watcher running you must read it faster than
`interval` or you will miss rounds.
## `GET /sync/read-status/message/{message_id}`
One message, one record — a point lookup rather than a batch:
```jsonc
{ "id": "AAMk…", "isRead": true,
"lastModifiedDateTime": "2026-08-07T10:14:52Z", "subject": "Invoice #421" }
```
Identical shape to an entry in a sync `value` list, so both parse with the same
code. What makes it different from the endpoints above:
- **Cursor-free.** Touches no delta cursor, no cached state, and advances
nothing. Call it as often as you like without affecting a sync in progress.
- **Live.** Reports the mailbox *now*, straight from Graph — not what the last
round happened to capture. That makes it the right tool for re-checking one
message ("has this been read yet?") and for confirming a status after the fact.
- **Any id.** Works whether or not the message appeared in a sync, and whatever
folder it lives in.
It costs one Graph call per message, so it's a lookup, not a substitute for
delta — walking a mailbox with it would be far slower than a single sync round.
```bash
curl -H "$A" "$B/sync/read-status/message/AAMkAGI2TG93..."
```
URL-encode the id. Ids containing `/`, `+`, or `=` are handled (the route uses a
`:path` converter), so an already-encoded `%2F` works too. Unknown or deleted
ids surface Graph's own `404 ErrorItemNotFound`.
## `GET /sync/read-status/status`
Health check for the whole subsystem.
| Field | Meaning |
| ----- | ------- |
| `watching` / `interval` | Is the poller thread alive, and at what period |
| `folder` | Folder the cursor belongs to |
| `last_sync_at` | Timestamp of the most recent round |
| `last_change_count` / `last_removed_count` | Size of that round |
| `has_delta_link` | A real cursor exists ⇒ running incrementally |
| `backfill_in_progress` | Paused mid-backfill at the page cap ⇒ more rounds to go |
| `last_error` | Last Graph failure from the background thread, else `null` |
`has_delta_link: false` + `backfill_in_progress: true` is the normal state
*during* a long first sync.
## `POST /sync/read-status/watch`
Starts a daemon thread that runs the same sync every `interval` seconds and
writes each change to stdout.
```jsonc
{ "interval": 60, "folder": "inbox" } // interval min 10, both optional
```
- Idempotent — a second POST while running just answers
`{"message": "Already watching read-status changes"}`.
- While a backfill is still incomplete the loop continues immediately instead of
sleeping out the interval, so a big first sync finishes in consecutive chunks.
- Delivery is `_emit_read_status_changes()`, which prints. **That's the hook
point** — replace it to push to Slack, a webhook, or a queue.
## `DELETE /sync/read-status/watch`
Signals the thread to stop; `404` if nothing is running. The cursor survives, so
restarting the watcher resumes from where it left off rather than re-backfilling.
---
## Typical first run
```bash
export EMAIL_API_TOKEN=...
A="Authorization: Bearer $EMAIL_API_TOKEN"
B=http://localhost:5000
curl -X POST -H "$A" $B/auth/start # sign in once (see README)
# Baseline. Keep calling while "complete": false.
curl -H "$A" "$B/sync/read-status?since=2026-08-01T00:00:00Z"
# From here on, each call returns only what changed.
curl -H "$A" "$B/sync/read-status"
# Or hand it to the background poller and read results out of /changes.
curl -X POST -H "$A" -H "Content-Type: application/json" \
-d '{"interval":60,"folder":"inbox"}' $B/sync/read-status/watch
curl -H "$A" $B/sync/read-status/status
curl -H "$A" $B/sync/read-status/changes
# Re-check one message any time — no cursor involved.
curl -H "$A" "$B/sync/read-status/message/AAMkAGI2TG93..."
```
## Which endpoint do I want?
| You want | Use |
| -------- | --- |
| Everything that changed since last time | `GET /sync/read-status` |
| The full list a round produced (or the watcher's) | `GET /sync/read-status/changes` |
| The status of **one** message you already have an id for | `GET /sync/read-status/message/{id}` |
| Continuous tracking without calling in a loop | `POST /sync/read-status/watch` |
| Whether any of the above is healthy | `GET /sync/read-status/status` |
Rule of thumb: **delta for "what changed", point lookup for "what about this
one".** Using the lookup in a loop over a mailbox works but costs one Graph call
per message — a single sync round does the same job in pages of 100.
## How the cursor works
- A finished round returns Graph's **deltaLink**, saved to
`.delta_cache.json` (override with `EMAIL_API_DELTA_CACHE`; in Docker it lives
on the `/data` volume beside the token cache). Keyed by user + folder — change
either and the cache is ignored rather than misapplied.
- A round stopped by `max_pages` has no deltaLink yet, so it saves Graph's
**nextLink** instead. That resume position takes priority over any older
deltaLink on the following call, which is what makes a capped backfill safe:
the cursor never advances past data you haven't received.
- Cursors expire. Graph answers `410 Gone`, and the sync automatically falls
back to a fresh baseline for that folder.
- `reset=true` throws the cursor away deliberately — expect a full backfill, and
pass `since` with it unless you want the whole history again.
## Limits worth knowing
- **Folder-scoped only.** `/me/messages/delta` is not supported by Graph. Watch
another folder by passing `folder=`, but each folder is its own cursor and the
disk cache holds one at a time — switching folders forces a re-backfill.
- **Polling, not push.** Latency floor is the poll `interval`. True push needs a
Graph change-notification subscription (public HTTPS endpoint, validation
handshake, ~3-day renewals) — and you'd keep delta anyway as the catch-up path
for dropped notifications.
- **Single worker.** Cursor, watcher thread, and the `last_changes` buffer are
in-memory per process, so this only behaves with one uvicorn worker (which is
what the Docker service runs, for the same reason auth needs it).