diff --git a/ARCHITECTURE-DETAIL.md b/ARCHITECTURE-DETAIL.md new file mode 100644 index 0000000..4903a63 --- /dev/null +++ b/ARCHITECTURE-DETAIL.md @@ -0,0 +1,469 @@ +# Utopia Pricing Agent β€” Architecture (engineering detail) + +The technical reference: data sources, decision cascade, formulas, file map, backtests. +For the plain-English overview see **[ARCHITECTURE.md](ARCHITECTURE.md)**. + +A one-page Streamlit dashboard (Utopia/CRAI design system) that turns **live COSMOS +data** into price recommendations a human can Approve / Modify / Reject. Read-only: +nothing is written back to COSMOS or Amazon. + +--- + +## 1. High-level view + +```mermaid +flowchart LR + U["πŸ§‘ User
(same-network browser)"] --> APP + + subgraph APP["app.py β€” presentation (Streamlit, CRAI theme)"] + SIDE["Sidebar
Single product / Product line
+ filters, kill switch"] + QUEUE["Recommendation queue
tiles Β· pills Β· rows"] + SECT["Per-SKU sections
price Β· inventory Β· scenarios
competitors Β· PPC Β· costs Β· AI"] + end + + subgraph DASH["dashboard/ package"] + THEME["theme.py
CRAI tokens + plotly template"] + LIVE["live_data.py
adapter + decision engine
+ scenario economics"] + end + + subgraph CORE["src/pricing_agent β€” analysis core"] + AN["analyze.py"] + MARGIN["margin_engine.py"] + ELAST["elasticity.py"] + PERF["performance.py"] + SVC["cosmos/service.py"] + CLIENT["cosmos/client.py"] + end + + COSMOS[("COSMOS API")] + APIFY[("Apify β€” optional")] + + APP --> THEME + APP --> LIVE + LIVE --> AN + AN --> MARGIN & ELAST & PERF + AN --> SVC --> CLIENT --> COSMOS + LIVE --> SVC + AN -.optional.-> APIFY +``` + +--- + +## 2. Layers + +| Layer | Files | Responsibility | +|---|---|---| +| **Presentation** | `app.py` | All rendering, zero pricing logic. Session state (approve/modify/reject, filters, per-SKU section + window), staged progress loader, session-state cache. | +| **Design system** | `dashboard/theme.py`, `.streamlit/config.toml` | CRAI palette (cream `#f4f0e8`, teal `#0c8276`, coral `#df4f33`, navy `#22304e`), Inter font, plotly template. | +| **Adapter + engine** | `dashboard/live_data.py` | Builds the per-SKU dict; **decides** the action; computes **scenario economics** (elasticity projection β†’ bulk reconciliation β†’ calibration); exposes `scenarios_for_window()`. | +| **Competitive state** | `src/pricing_agent/competitive_state.py` | The one competitive fact the cascade may read. Adapts either scraper into a typed `WON`/`LOST_PRICE`/`LOST_ELIGIBILITY`/`SUPPRESSED` state with a source and a timestamp, gates it on age, and logs disagreement between sources. | +| **Analysis core** | `src/pricing_agent/analyze.py` | Orchestrates one SKU: fees β†’ trend β†’ bulk β†’ ad cost β†’ elasticity β†’ actual-profit evidence. | +| **Money math** | `tools/margin_engine.py` | Pure: break-even, MAP, contribution margin, suggested price. | +| **Statistics** | `elasticity.py`, `performance.py` | Log-log elasticity fit, profit-optimal sweep, actual-profit aggregation. | +| **Data access** | `cosmos/{client,service,models}.py` | Auth + retry client; endpoint calls + response flattening; typed pydantic models. | + +--- + +## 3. Data sources β€” what each COSMOS endpoint feeds + +```mermaid +flowchart TB + subgraph COSMOS["COSMOS API"] + TH["/sales-insight/takehome-calculator
nested fees.breakdown Β· cost.breakdown"] + INVP["/invp-insight
trend + inventory + dateMap PROJECTIONS"] + BULK["/sales-insight/bulk-calculator
storage + total take-home"] + SI["/sales-insight (daily, 6-month)
price Β· units Β· revenue Β· profit Β· ad spend"] + PROD["/products
brand Β· marketplace"] + CAMP["/api/campaigns Β· /adsApi
budget Β· ACoS Β· ad sales (not yet wired)"] + end + + TH -->|"_flatten_takehome()"| FEES["Fee model
referral% Β· FBA Β· landed Β· returns"] + INVP --> TREND["Velocity + cover days"] + INVP --> INVPROJ["Inventory Outlook tab
real weekly units/value/cover/arrivals"] + BULK --> STORAGE["Storage + take-home (scenarios)"] + SI --> HIST["180-day daily series
(window filter + calibration)"] + SI --> ADS["Ad spend / TACoS (PPC tab)"] + PROD --> META["Brand / marketplace"] +``` + +**Three response quirks handled:** + +- **Fees come nested** (`fees.breakdown["Referral Fee"]`, `"$ 9.28"` strings). + `service._flatten_takehome()` normalises them β€” without it every fee parsed to 0 (the + old "$0.99 / break-even $0" bug). +- **INVP `dateMap`** holds COSMOS's own **forward inventory projection** (weekly units, + value, cover days, warehouse arrivals). The Inventory tab renders this directly β€” not + a locally-invented forecast. +- **SKU lookups are CONTAINS/relevance searches, not equality.** See Β§3.1 β€” this one + silently bound the wrong product's data to a SKU. + +### 3.1 Exact-SKU joins (`service._exact_row`) + +Neither `/api/products` nor `/api/invp-insight` has an equality filter, and both used to +fall back to `data[0]` "if COSMOS returned a fuzzy match set". That bound one product's +ASIN, cost, brand and inventory to a **different product's SKU**. + +Observed live: `get_product("UBCFKFITTEDSHEETWHITECALKING")` β€” a SKU COSMOS does not carry +at all β€” returned `UBMICROFIBERGUSSETPILLOWWHITEQUEEN` / `B08DTH86Q2`. + +Two independent problems, and each guard is necessary: + +1. **Wrong parameter.** `get_product` queried `q=` (relevance across the whole catalogue). + For `UBMICROFIBERDUVETTWINWHITE` the correct row sat on **page 2 of 100-row pages**, + behind 100 unrelated products β€” a 20-row lookup never saw it. `sku=` returns it first. +2. **`sku=` is still a CONTAINS filter.** `sku=UBCFKMATTRESSPROTECTORTWIN88` returns three + rows: the real one (`B00MRH9NCK`), the `...BOX` variant (`B09K7HXJ4M`), and a + `WAL...` Walmart row whose "ASIN" (`8946709597`) is not an ASIN. **All three are + `marketplace: AMAZON_USA`**, so the marketplace check alone does not separate them β€” + only the exact SKU test does. + +`_exact_row()` requires an exact SKU match **and** the right marketplace, and returns +`None` otherwise, logging what it rejected. Blast radius of the old behaviour: +`analyze_price` assigns `asin = product.asin` whenever INVP has none, so the competitive +scrape would have run against an unrelated listing; and `get_invp`'s `skuPrefix` matches +every colour variant, so a sibling's inventory and cover days fed `LOW_STOCK` / +`EXCESS_STOCK` directly β€” a sibling reading 12 units / 3 days would fire a stockout raise +on a SKU holding 4,000 units. + +Genuinely **not** in this COSMOS integration β†’ shown as "β€”", never faked: ad-attributed +sales / ACoS / campaign budget (live in `/api/campaigns` + `/adsApi`, not yet wired), and +historical competitor prices (Apify gives a current snapshot only). + +**Competitor data is no longer display-only.** COSMOS has no Buy Box, no rival price and no +third-party offer anywhere in it β€” that gap is filled by a scrape, and two of its facts +(our Buy Box being suppressed, and a rival materially undercutting us) reach the verdict. +They are the only two, they are bounded by the guardrails, and their absence changes +nothing. See Β§6. + +--- + +## 4. Per-SKU pipeline (one "Analyze") + +```mermaid +sequenceDiagram + participant U as User + participant A as app.py + participant L as live_data.build_live_sku + participant AN as analyze.py + participant S as CosmosService + U->>A: Single product / Product line + A->>L: get_live_data(skus, progress_cb) + Note over A: staged progress bar (3β†’10β†’45β†’82β†’94β†’100%) + L->>S: get_current_price + L->>S: get_sales_history (180d, parallel windows) + L->>AN: analyze_price (fees, trend, bulk, elasticity, evidence) + L->>S: get_invp (real inventory projection) + L->>S: bulk_quote (storage) + L->>L: decide action + scenario economics + 30d/6mo calibration + L-->>A: {summary, details, errors} (session-cached) + A-->>U: queue + expandable per-SKU analysis +``` + +--- + +## 5. Scenario economics (the heart of the Scenarios tab) + +For each candidate price, one consistent chain: + +```mermaid +flowchart LR + W["Window filter
7/14/30/90d Β· 6mo"] --> BASE["Baseline velocity
= avg units/day in window"] + BASE --> DEMAND["Units(p) = units Γ— (p/cur)^elasticity"] + DEMAND --> REV["Revenue = units Γ— p Γ— 30"] + REV --> AD["Ad spend = TACoS Γ— revenue"] + DEMAND --> TH["Take-home (bulk calculator fee model)"] + TH --> GROSS["Gross = take-home βˆ’ storage βˆ’ ad"] + GROSS --> CAL["Γ— realization factor
(actual booked Γ· modeled at current)"] + CAL --> NET["Net profit / 30d"] +``` + +Key rules: + +- **Current row = FACT**, not a projection: real units, real revenue, real ad spend, real + booked profit. Its price is the **average price sold** (revenue Γ· units) so + `price Γ— units Γ— 30 = revenue` reconciles β€” this is *below* list when promos ran, and + changes with the window because the avg selling price differed period to period. The + **list price is fixed**. +- **Calibration:** raw bulk-calculator profit over-states reality (prices at list, ignores + real returns/promos). A **realization factor** = actual booked profit Γ· modeled profit at + the current price scales every projected row. +- **The elasticity used for projections is gated** (`projection_elasticity()`). It honours + the same `actionable` test the decision does: a slope whose 95% CI spans zero β€” or a + **positive** slope, which `estimate_elasticity` can return since `actionable` requires + `e < 0` but the value is not clamped β€” falls back to `FALLBACK_ELASTICITY`. Ungated, a + positive slope projected that *raising* price sells *more*, and that number drives the + 30-day impact tile, the portfolio opportunity total and the queue sort. + `elasticity_is_fitted` records which was used. +- **⭐** marks the highest-net-profit price in the current view. + +--- + +## 6. Decision engine (deterministic, first match wins) + +```mermaid +flowchart TD + S([signals]) --> R0{cost data = 0?} + R0 -- yes --> INV["πŸ” INVESTIGATE Β· NO_COST_DATA"] + R0 -- no --> RB{our Buy Box suppressed?} + RB -- yes --> INVB["πŸ” INVESTIGATE Β· BUYBOX_SUPPRESSED"] + RB -- no --> R1{price < break-even?} + R1 -- yes --> UP1["↑ raise to safe floor Β· BELOW_BREAK_EVEN"] + R1 -- no --> R2{losing money after ads?} + R2 -- yes --> UP2["↑ raise Β· LOSING_MONEY"] + R2 -- no --> R3{cover ≀ 35d?} + R3 -- yes --> UP3["↑ +5% Β· LOW_STOCK"] + R3 -- no --> R4{βˆ’30% sales, no cause?} + R4 -- yes --> INV2["πŸ” INVESTIGATE Β· UNEXPLAINED_DROP"] + R4 -- no --> R5{cover β‰₯ 90d?} + R5 -- yes --> DN["↓ βˆ’5% Β· EXCESS_STOCK"] + R5 -- no --> RC{rival β‰₯3% below AND corroborated?} + RC -- yes --> DNC["↓ toward rival Β· COMPETITOR_UNDERCUT"] + RC -- no --> R6{profit-optimal β‰  current?} + R6 -- yes --> MOVE["↑/↓ toward optimal Β· PROFIT_OPTIMAL"] + R6 -- no --> R7{a price we ran booked more?} + R7 -- yes --> BO["↑/↓ toward it Β· BEST_OBSERVED"] + R7 -- no --> HOLD["β†’ MAINTAIN Β· NO_SIGNALS"] +``` + +Guardrails: floor = highest of four break-evens Γ— 1.02, ceiling = current Γ— 1.25; the +recommended move is capped at Β±5% per step; the kill switch pauses all approvals. + +### AD_SPIRAL β€” applied after the cascade + +When ad cost per unit climbs almost as fast as price, each extra $1 of price buys only +cents of contribution and **no** price reaches break-even. That verdict (`AD_SPIRAL`, +Investigate-and-hold) is applied *after* the cascade and overrides whatever fired β€” so +what it must **not** override is named explicitly: + +```python +AD_SPIRAL_YIELDS_TO = frozenset({"NO_COST_DATA", "BUYBOX_SUPPRESSED", "LOW_STOCK"}) +``` + +| Yields to | Why | +|---|---| +| `NO_COST_DATA` | With `costPerUnit`/`fbaFee` at 0, `fixed` is understated, which makes the `contribution <= 0` test **easier** to hit. A missing-COGS SKU would be sent to the ad console when the fix is a data-entry field. | +| `BUYBOX_SUPPRESSED` | A listing nobody can buy from has no meaningful ad economics. | +| `LOW_STOCK` | The only one that changes a **price**, not just a label. A shelf about to empty gets +5% whatever the ad slope does β€” those units sell regardless, so the only question is what we get for them. Holding sells the last of the stock cheap. | + +Deliberately narrow: `EXCESS_STOCK` does **not** outrank it, because cutting price to clear +stock is exactly the move that cannot work when ads eat the contribution. + +### Competitor rules β€” the two that can move a price, and what bounds them + +Both branches read a single `CompetitiveState` +([competitive_state.py](src/pricing_agent/competitive_state.py)), never a raw scrape: + +| Rule | Fires when | Effect | +|---|---|---| +| `BUYBOX_SUPPRESSED` | Amazon is not showing our offer | **Investigate, hold.** Placed directly under `NO_COST_DATA`: the only two states where the answer is "go and fix something" rather than "set a price". | +| `COMPETITOR_UNDERCUT` | cheapest rival β‰₯ `competitor_undercut_material_pct` below us, **and the basis qualifies** (below) | **Decrease toward the rival**, floored and step-capped like every other branch. | + +**Basis decides what "qualifies" means**, because `competitor_min` can be a price of two +different things: + +- `same-asin-buybox` (Apify) β€” another seller's offer on **our own listing**. Only + `LOST_PRICE` fires; a rival holding the Buy Box *above* us is `LOST_ELIGIBILITY`, where + cutting donates margin. Losing the Buy Box on price **is** the corroboration. +- `like-for-like-sheet` β€” a **rival brand's** equivalent variant, matched on size + colour. + This additionally requires **corroboration**: either demand has materially dropped, or we + are not actually winning the Buy Box. + + *Why:* on its own a sheet row says only "a different brand is cheaper". True, reportable, + but not evidence the gap is costing us anything β€” we can sit 3% dearer, hold our own Buy + Box and sell fine on brand, reviews or the Prime badge. And because this branch sits + **above** `PROFIT_OPTIMAL`, an uncorroborated cut could overrule an elasticity fit that + wanted a *raise*. Switch: `competitor_sheet_requires_corroboration` (default `true`). + + The documented `UNEXPLAINED_DROP` fall-through is unaffected β€” + `UBMICROFIBERDUVETKINGPURPLE` and `UBMICROFIBERBS4PCFULLGREY` arrive here *with* a + velocity drop, which is the first form of corroboration. + + An uncorroborated material undercut is still **reported** in the root cause ("Competitor + undercut not acted on"), so it never looks like missing data. + +A third signal, a **competitor premium** while we hold the Buy Box, is a narrative note +only. It never sets a price and never changes an action. + +Ordering is deliberate: **inventory risk outranks competitor position, which outranks +profit-optimal.** Chasing a rival down while the shelf is emptying pays margin to sell out +faster. + +Three properties make this safe to ship: + +1. **Fail-safe.** Absent, failed, stale (> `competitor_state_max_age_hours`) and + "ownership unknown" all collapse to one flag, and the cascade then computes exactly the + verdict it computed before competitor data existed. Competitor data can only ever *add* + a verdict. +2. **Never below break-even.** The rival price is a *candidate* (`comp_match`), not a + decision. +3. **One named reason per verdict.** No blended scores β€” every fired rule is traceable to a + single reason code, and `logger.info` names the SKU, the rule, the state and the source. + +Thresholds live in [config/pricing_rules.yaml](config/pricing_rules.yaml) +(`competitor_undercut_material_pct: 0.03`, `competitor_premium_material_pct: 0.10`, +`competitor_state_max_age_hours: 6.0`, `competitor_sheet_requires_corroboration: true`), +not in code. The 3% floor sits above the ~2% band our own realized price already swings +through as coupons toggle. + +### Inventory cover matches COSMOS Inventory Planning + +`cover_days` **is COSMOS's own `coverDays`**, so the dashboard and the INVP grid never quote +two different numbers for one SKU. COSMOS counts **inbound** stock against a **7-day** +velocity, so it reads longer than what is on the shelf β€” `UBMICROFIBERDUVETTWINWHITE` is +78 days on (3,999 on hand + 1,030 inbound) Γ· 64/day, against 63 on-hand-only. Both are +reported: the tile leads with the matched figure and appends `63 d on hand, rest inbound`. + +The on-hand figure remains the **fallback**, because COSMOS returns `coverDays: 0` on some +very low-velocity SKUs that hold months of stock (`UBMICROFIBERBS4PCFULLGREY`: 167 units, +334 real days, COSMOS said `0`). Zero satisfies neither inventory rule, so taken literally it +silences both. The reconciled figure is computed **before** `_decide` and passed in, so the +tile and the rule that fired read the same number by construction. + +**Trade-off, accepted deliberately:** stockout risk is now judged partly on stock that has not +landed. Measured over the 56-SKU covered line, matching COSMOS moved 7 verdicts β€” +`LOW_STOCK` 7 β†’ 4, `EXCESS_STOCK` 18 β†’ 22. The one to watch is +`UBMICROFIBERBS4PCKINGWHITE`: **12 days on the shelf, 84 with inbound**, so it no longer +raises. If that shipment slips, nothing protects it. + +Display bands are COSMOS's Alpha/Beta scheme (`theme.COVER_BANDS`, Alpha 20/40/70/100). The +pricing **triggers** are separate and live in `pricing_rules.yaml` +(`low_cover_days: 35`, `high_cover_days: 90`) β€” COSMOS's pink at 70 days is a *replenishment* +warning, while crossing a trigger here spends margin on a 5% move. + +### Coverage: the comparison sheet gates competitor data, one product line at a time + +The competitor workbook currently covers **one product line**, so the engine reads it as the +first competitor source and **gates on coverage**: + +| SKU | Competitor state | +|---|---| +| In the sheet | Priced from the sheet β€” real like-for-like rival prices, `basis=like-for-like-sheet` | +| Not in the sheet | **`N/A`**, naming what the sheet *does* cover. No rule fires; the verdict is byte-identical to the competitor-blind one | + +**Coverage is the exact SKU set in the sheet, not a line prefix.** Measured against the real +workbook, a prefix gate would be wrong in both directions: the `UBMICROFIBERDUVET` run contains +49 `UBMICROFIBERDUVET*` SKUs **and 7 `UBMICROFIBERBS4PC*`** ones, while the line has 139 SKUs +in COSMOS of which only 56 reached a comparison row. So the sheet's own SKU list is the +authority, and "not in the sheet" is a **coverage hole**, never a claim that the SKU has no +competitors. + +`competitor_sheet_only: true` (default while one line is under test) means an uncovered SKU gets +N/A rather than falling through to a per-ASIN Apify scrape. Config: `competitor_sheet_path` +(blank = auto-discover the newest `Competitor_Price_Comparison_*.xlsx`), `competitor_sheet_dirs`, +`competitor_sheet_max_age_hours: 168` (the sheet is a 25–35 min batch run, not a live feed). + +Two sheet-driven refinements, both from real rows: + +- **A rival whose own Buy Box is suppressed is excluded from the band.** Their price is not + buyable, so undercutting it donates margin for nothing. +- **A material undercut can explain a velocity drop** (subject to the corroboration rule + above), converting an `UNEXPLAINED_DROP` Investigate into an actionable verdict β€” exactly + as the existing stockout branch already did. + +### Two scrapers, one state + +| Source | Authoritative for | Why | +|---|---|---| +| **Apify** (`tools/amazon/apify.py`) | **Buy Box state read by the engine** | The only source carrying a seller id, so the only one that can tell `WON` from `LOST_PRICE` from `LOST_ELIGIBILITY`. Those lead to opposite actions. | +| **Playwright** (`../scraper/`) | The workbook: like-for-like size/colour matching, BSR, demand buckets, SKU gaps | Apify cannot produce any of it. Its Buy Box field knows only whether a price *rendered*, not whose it was. | + +The split is by **question**, not preference. `reconcile()` cross-checks the authoritative +state against the Playwright run's own cache (`scraper/.scrape_cache.json`, keyed `ASIN@ZIP`) +and **logs any disagreement**. A disagreement never changes the verdict. + +--- + +## 6.1 Verdict backtest + +`scripts/backtest_competitor_rules.py` re-runs the **captured real cascade inputs** (real +`AnalysisResult`, real 180-day history, real fee stack, real scenario grid) through the same +`_decide`, varying only the arm. Nothing is reconstructed by hand. + +Competitor rules blind vs live (7 requested SKUs, 6 analysed β€” see Β§9): + +| Arm | Verdicts changed | +|---|---| +| Real competitor state | **0 / 6** β€” the fail-safe working | +| Counterfactual 8% undercut | 4 / 6 β€” the others blocked by `LOW_STOCK` / `LOSING_MONEY` | +| Counterfactual suppression | 6 / 6 β†’ Investigate | + +Policy delta, old vs new (the two cascade changes), same SKUs: + +| Case | Changed | Notes | +|---|---|---| +| `SHEET_UNDERCUT_WE_WIN` | 2 / 6 | Both were being **cut while holding the Buy Box with flat demand**: `UBMICROFIBERDUVETTWINWHITE` Decrease $17.06 β†’ **Increase $17.94** (`BEST_OBSERVED`), `UBMICROFIBERGUSSETPILLOWWHITEQUEEN` Decrease $25.01 β†’ **Increase $27.37**. The two corroborated SKUs were unchanged. | +| `AD_SPIRAL_LOW_STOCK` | 4 / 6 | `Investigate/AD_SPIRAL` β†’ `Increase/LOW_STOCK` | +| `AD_SPIRAL_NO_COST` | 6 / 6 | `Investigate/AD_SPIRAL` β†’ `Investigate/NO_COST_DATA` | +| **Targets below break-even, any arm** | **0** | | + +**The invariant is measured on the TARGET, not the step-capped first move.** A SKU already +selling under its own ad-inclusive floor cannot be lifted over it in one 5% step, and +reporting that deliberate multi-cycle climb as a breach buries any real one. Two SKUs are +below floor and stepping up by design; both are reported separately. + +--- + +## 7. Key formulas + +| Quantity | Formula | +|---|---| +| Take-home / unit | `pΒ·(1 βˆ’ referral% βˆ’ returns%) βˆ’ landed βˆ’ FBA βˆ’ other` | +| Break-even | `(landed + FBA + returns + other) / (1 βˆ’ referral%)` | +| Elasticity | OLS on `ln(units/day) = a + eΒ·ln(price)` over 6 months | +| Scenario demand | `units Γ— (p / pβ‚€)^e` | +| Realization factor | `actual booked profit (window) Γ· modeled net at current price` | +| TACoS | `ad spend Γ· total revenue` (window) | +| Avg sold price | `revenue Γ· units` (window) β€” reconciles the Current row | +| Best observed price | `avg_price` of the best-earning band β€” **the price actually charged**, never the $0.50-rounded `price_band` key | + +--- + +## 8. Repository map + +``` +pricing_agent/ +β”œβ”€β”€ app.py # dashboard (presentation only) +β”œβ”€β”€ legacy_app.py # previous analyst UI (still runnable) +β”œβ”€β”€ dashboard/ +β”‚ β”œβ”€β”€ theme.py # CRAI design tokens + plotly template +β”‚ └── live_data.py # COSMOS adapter, decision + scenario engine +β”œβ”€β”€ src/pricing_agent/ +β”‚ β”œβ”€β”€ analyze.py # per-SKU orchestration β†’ AnalysisResult +β”‚ β”œβ”€β”€ competitive_state.py # canonical Buy Box state + two-scraper reconciliation +β”‚ β”œβ”€β”€ elasticity.py # demand model + profit optimizer +β”‚ β”œβ”€β”€ performance.py # actual-profit evidence +β”‚ β”œβ”€β”€ tools/margin_engine.py # pure fee/break-even math (golden-tested) +β”‚ └── cosmos/ +β”‚ β”œβ”€β”€ client.py # auth + retry HTTP +β”‚ β”œβ”€β”€ service.py # endpoints, _exact_row, _flatten_takehome, INVP projections +β”‚ └── models.py # typed COSMOS responses (pydantic) +β”œβ”€β”€ config/ # settings + pricing_rules.yaml (incl. competitor thresholds) +β”œβ”€β”€ scripts/ +β”‚ └── backtest_competitor_rules.py # verdict delta + old-vs-new policy delta +β”œβ”€β”€ .streamlit/config.toml # CRAI theme +└── tests/ # margin-engine golden values, test_competitor_rules.py + # (fail-safe + ordering invariants), test_pricing_safety.py, + # test_cosmos.py (exact-SKU joins) +``` + +--- + +## 9. Known data gaps + +- **Ad-attributed sales / ACoS / campaign budget** β€” live in `/api/campaigns` + `/adsApi`, + not yet wired. Shown as "β€”". +- **Historical competitor prices** β€” Apify returns a current snapshot only. +- **`UBCFKFITTEDSHEETWHITECALKING` is not in COSMOS at all.** `sku=` returns zero rows on + any marketplace and the fee endpoint replies "Product not found". Likely delisted or + renamed. It now fails loudly (`get_product` β†’ `None`) rather than adopting another + product's identity. A catalogue question, not a code one. + +--- + +## 10. Principles + +1. **Deterministic core, narrative shell** β€” every number is a formula over COSMOS data; language models only phrase explanations. +2. **Read-only** β€” the agent proposes; a human approves; nothing writes back. `submit_price_approval` remains a stub with no callers. +3. **Honest gaps** β€” missing upstream data shows "β€”" or an explicit investigation, never a fabricated number. A wrong number is worse than a blank one. +4. **Facts vs projections are labeled** β€” the Current row is real booked history; other prices are clearly modeled. +5. **Everything reconciles** β€” one averaging window drives units, revenue, ads and profit so `price Γ— units = revenue` always holds. diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 6d7bdaf..4cc8c5d 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -1,378 +1,187 @@ -# Utopia Pricing Agent β€” Architecture +# How the Pricing Agent Works -A one-page Streamlit dashboard (Utopia/CRAI design system) that turns **live COSMOS -data** into price recommendations a human can Approve / Modify / Reject. Read-only: -nothing is written back to COSMOS or Amazon. +A plain-English guide. No coding or Amazon knowledge needed. +> Engineers: the detailed reference β€” data sources, formulas, file map, backtests β€” +> now lives in **[ARCHITECTURE-DETAIL.md](ARCHITECTURE-DETAIL.md)**. --- -## 1. High-level view +## What it is -```mermaid -flowchart LR - U["πŸ§‘ User
(same-network browser)"] --> APP +We sell products on Amazon. Every product has a price, and the right price is a +constant question: too high and we stop selling, too low and we lose money on every +sale. - subgraph APP["app.py β€” presentation (Streamlit, CRAI theme)"] - SIDE["Sidebar
Single product / Product line
+ filters, kill switch"] - QUEUE["Recommendation queue
tiles Β· pills Β· rows"] - SECT["Per-SKU sections
price Β· inventory Β· scenarios
competitors Β· PPC Β· costs Β· AI"] - end +This tool looks at one product at a time, works out what its price *should* be, and +shows a person the answer with the reasoning behind it. - subgraph DASH["dashboard/ package"] - THEME["theme.py
CRAI tokens + plotly template"] - LIVE["live_data.py
adapter + decision engine
+ scenario economics"] - end - - subgraph CORE["src/pricing_agent β€” analysis core"] - AN["analyze.py"] - MARGIN["margin_engine.py"] - ELAST["elasticity.py"] - PERF["performance.py"] - SVC["cosmos/service.py"] - CLIENT["cosmos/client.py"] - end - - COSMOS[("COSMOS API")] - APIFY[("Apify β€” optional")] - - APP --> THEME - APP --> LIVE - LIVE --> AN - AN --> MARGIN & ELAST & PERF - AN --> SVC --> CLIENT --> COSMOS - LIVE --> SVC - AN -.optional.-> APIFY -``` +**It never changes a price.** It makes a recommendation. A human reads it and clicks +Approve, Modify, or Reject. Nothing is sent to Amazon. --- -## 2. Layers +## The one-sentence version -| Layer | Files | Responsibility | +> It reads our real sales history, works out what each sale actually earns after all +> costs, checks whether we have too much or too little stock, looks at what rivals +> charge, and then suggests a price β€” showing its working, and refusing to suggest +> anything that would lose money. + +--- + +## Where the numbers come from + +Everything starts from **COSMOS**, our internal system that already holds our Amazon +data. The tool reads five things from it: + +| What | Why it matters | +|---|---| +| **Fees and costs** | What Amazon charges us per sale, plus what the product cost us | +| **Sales history** (6 months, daily) | What we charged, how many we sold, what we actually earned | +| **Stock levels** | How many units we have, and how long they'll last | +| **Storage charges** | What Amazon bills us to warehouse unsold stock | +| **Advertising spend** | What we paid in ads to make those sales | + +**What the 6 months of history is for:** it tells us which prices we have *already* +tried and what each one really earned β€” so the tool can recommend a price we have +proof about, instead of guessing. It also spots when sales have dropped well below +normal, which is a signal something has changed. + +**What competitor prices are for:** only two things. If Amazon has hidden our listing, +stop and fix that. If a rival is meaningfully cheaper **and** it is visibly costing us +sales, move toward their price β€” but never below the point where we lose money. + +--- + +## How it decides + +The tool runs down a checklist, **in order, and stops at the first thing that +applies.** This is deliberate β€” it means every recommendation traces back to exactly +one reason, and you can always ask "why this price?" and get a single answer. + +Roughly top to bottom: + +| # | If we find this… | …the answer is | |---|---|---| -| **Presentation** | `app.py` | All rendering, zero pricing logic. Session state (approve/modify/reject, filters, per-SKU section + window), staged progress loader, session-state cache. | -| **Design system** | `dashboard/theme.py`, `.streamlit/config.toml` | CRAI palette (cream `#f4f0e8`, teal `#0c8276`, coral `#df4f33`, navy `#22304e`), Inter font, plotly template. | -| **Adapter + engine** | `dashboard/live_data.py` | Builds the per-SKU dict; **decides** the action; computes **scenario economics** (elasticity projection β†’ bulk reconciliation β†’ calibration); exposes `scenarios_for_window()`. | -| **Competitive state** | `src/pricing_agent/competitive_state.py` | The one competitive fact the cascade may read. Adapts either scraper into a typed `WON`/`LOST_PRICE`/`LOST_ELIGIBILITY`/`SUPPRESSED` state with a source and a timestamp, gates it on age, and logs disagreement between sources. | -| **Analysis core** | `src/pricing_agent/analyze.py` | Orchestrates one SKU: fees β†’ trend β†’ bulk β†’ ad cost β†’ elasticity β†’ actual-profit evidence. | -| **Money math** | `tools/margin_engine.py` | Pure: break-even, MAP, contribution margin, suggested price. | -| **Statistics** | `elasticity.py`, `performance.py` | Log-log elasticity fit, profit-optimal sweep, actual-profit aggregation. | -| **Data access** | `cosmos/{client,service,models}.py` | Auth + retry client; endpoint calls + response flattening; typed pydantic models. | +| 1 | Cost information is missing | **Stop.** Can't price a product without knowing what it cost. Go fill it in. | +| 2 | Amazon isn't showing our listing to buyers | **Stop.** A hidden listing sells nothing at any price. Fix the listing first. | +| 3 | We're selling below what it costs us | **Raise the price.** Every sale is losing money. | +| 4 | Ads are eating more than the sale earns | **Raise the price.** | +| 5 | We're about to run out of stock | **Raise a little (5%)** to slow sales until more arrives. | +| 6 | Sales dropped sharply and we don't know why | **Stop and investigate.** Don't guess with the price. | +| 7 | We have far too much stock sitting there | **Lower a little (5%)** to shift it before storage costs mount. | +| 8 | A rival is meaningfully cheaper *and* it's actually costing us sales | **Lower toward their price** β€” but never below our own break-even. | +| 9 | The sales data suggests a more profitable price | **Move toward it.** | +| 10 | None of the above | **Leave it alone.** | + +**Notice that stock problems are ranked above competitor problems.** Cutting price to +chase a rival while the shelf is emptying just means selling out faster for less +money. --- -## 3. Data sources β€” what each COSMOS endpoint feeds +## The safety rails -```mermaid -flowchart TB - subgraph COSMOS["COSMOS API"] - TH["/sales-insight/takehome-calculator
nested fees.breakdown Β· cost.breakdown"] - INVP["/invp-insight
trend + inventory + dateMap PROJECTIONS"] - BULK["/sales-insight/bulk-calculator
storage + total take-home"] - SI["/sales-insight (daily, 6-month)
price Β· units Β· revenue Β· profit Β· ad spend"] - PROD["/products
brand Β· marketplace"] - CAMP["/api/campaigns Β· /adsApi
budget Β· ACoS Β· ad sales (not yet wired)"] - end +These apply to every recommendation, no exceptions: - TH -->|"_flatten_takehome()"| FEES["Fee model
referral% Β· FBA Β· landed Β· returns"] - INVP --> TREND["Velocity + cover days"] - INVP --> INVPROJ["Inventory Outlook tab
real weekly units/value/cover/arrivals"] - BULK --> STORAGE["Storage + take-home (scenarios)"] - SI --> HIST["180-day daily series
(window filter + calibration)"] - SI --> ADS["Ad spend / TACoS (PPC tab)"] - PROD --> META["Brand / marketplace"] -``` - -**Two response quirks handled:** -- **Fees come nested** (`fees.breakdown["Referral Fee"]`, `"$ 9.28"` strings). `service._flatten_takehome()` normalises them β€” without it every fee parsed to 0 (the old "$0.99 / break-even $0" bug). -- **INVP `dateMap`** holds COSMOS's own **forward inventory projection** (weekly units, value, cover days, warehouse arrivals). The Inventory tab renders this directly β€” not a locally-invented forecast. - -Genuinely **not** in this COSMOS integration β†’ shown as "β€”", never faked: ad-attributed -sales / ACoS / campaign budget (live in `/api/campaigns` + `/adsApi`, not yet wired), and -historical competitor prices (Apify gives a current snapshot only). - -**Competitor data is no longer display-only.** COSMOS has no Buy Box, no rival price and no -third-party offer anywhere in it β€” that gap is filled by a scrape, and as of the competitor -wiring two of its facts (our Buy Box being suppressed, and a rival materially undercutting us) -reach the verdict. They are the only two, they are bounded by the guardrails, and their -absence changes nothing. See Β§6. +- **Never below break-even.** However cheap a rival is, the tool will not suggest a + price that loses money. +- **Never more than 5% at once.** Big price jumps confuse both customers and our own + measurements. Large moves happen over several steps, each one checked. +- **Never above 25% up from today**, so a modelling error can't produce an absurd + price. +- **A kill switch** in the sidebar pauses all approvals instantly. --- -## 4. Per-SKU pipeline (one "Analyze") +## Two things it is careful about -```mermaid -sequenceDiagram - participant U as User - participant A as app.py - participant L as live_data.build_live_sku - participant AN as analyze.py - participant S as CosmosService - U->>A: Single product / Product line - A->>L: get_live_data(skus, progress_cb) - Note over A: staged progress bar (3β†’10β†’45β†’82β†’94β†’100%) - L->>S: get_current_price - L->>S: get_sales_history (180d, parallel windows) - L->>AN: analyze_price (fees, trend, bulk, elasticity, evidence) - L->>S: get_invp (real inventory projection) - L->>S: bulk_quote (storage) - L->>L: decide action + scenario economics + 30d/6mo calibration - L-->>A: {summary, details, errors} (session-cached) - A-->>U: queue + expandable per-SKU analysis -``` +**A blank is better than a wrong number.** If some piece of data is missing, the tool +shows a dash and says why. It never quietly fills in a zero or a guess β€” a wrong +number that looks confident is far more dangerous than an obvious gap. + +**It separates what happened from what it predicts.** Rows showing real past results +are labelled as facts. Rows showing "if we priced at X" are labelled as estimates. The +tool also reports how wrong it has been on that specific product in the past, so you +know how much to trust the estimate. --- -## 5. Scenario economics (the heart of the Scenarios tab) +## What it deliberately does not do -For each candidate price, one consistent chain: - -```mermaid -flowchart LR - W["Window filter
7/14/30/90d Β· 6mo"] --> BASE["Baseline velocity
= avg units/day in window"] - BASE --> DEMAND["Units(p) = units Γ— (p/cur)^elasticity"] - DEMAND --> REV["Revenue = units Γ— p Γ— 30"] - REV --> AD["Ad spend = TACoS Γ— revenue"] - DEMAND --> TH["Take-home (bulk calculator fee model)"] - TH --> GROSS["Gross = take-home βˆ’ storage βˆ’ ad"] - GROSS --> CAL["Γ— realization factor
(actual booked Γ· modeled at current)"] - CAL --> NET["Net profit / 30d"] -``` - -Key rules: -- **Current row = FACT**, not a projection: real units, real revenue, real ad spend, real - booked profit. Its price is the **average price sold** (revenue Γ· units) so - `price Γ— units Γ— 30 = revenue` reconciles β€” this is *below* list when promos ran, and - changes with the window because the avg selling price differed period to period. The - **list price is fixed**. -- **Calibration:** raw bulk-calculator profit over-states reality (prices at list, ignores - real returns/promos). A **realization factor** = actual booked profit Γ· modeled profit at - the current price scales every projected row, anchoring net profit to what the SKU truly - earns. -- **Suggested price** (teal callout) is computed from the **full 6-month** window always β€” - stable β€” independent of the display-window filter. -- **⭐** marks the highest-net-profit price in the current view. +- **It does not set prices.** Advisory only. Every change is a human decision. +- **It does not let the AI decide anything.** An AI writes the plain-English summary + you read β€” but every number and every recommendation comes from fixed arithmetic. + The AI explains; it never calculates or chooses. +- **It does not guess at missing data.** --- -## 6. Decision engine (deterministic, first match wins) +## Words you'll see + +| Term | Plain meaning | +|---|---| +| **SKU** | Our internal code for one specific product β€” e.g. a queen duvet in white | +| **ASIN** | Amazon's code for the same thing | +| **Buy Box** | The "Add to Cart" button. Several sellers can offer the same item; Amazon picks one to be the default. Win it and you get nearly all the sales. Lose it and sales collapse β€” so this matters enormously. | +| **Suppressed** | Amazon has hidden our listing entirely. Nobody can buy it. | +| **Break-even** | The price where we make exactly zero. Below it, every sale loses money. | +| **Margin** | What's left over from a sale after every cost | +| **Cover days** | How many days our current stock will last at the rate we're selling | +| **Elasticity** | How much sales volume changes when price changes. Some products lose lots of sales from a small rise; others barely notice. | +| **TACoS** | Advertising spend as a share of sales revenue | +| **Backtest** | Checking a method against past data to see how accurate it would have been | + +--- + +## The whole flow, start to finish ```mermaid flowchart TD - S([signals]) --> R0{cost data = 0?} - R0 -- yes --> INV["πŸ” INVESTIGATE Β· NO_COST_DATA"] - R0 -- no --> RB{our Buy Box suppressed?} - RB -- yes --> INVB["πŸ” INVESTIGATE Β· BUYBOX_SUPPRESSED"] - RB -- no --> R1{price < break-even?} - R1 -- yes --> UP1["↑ raise to safe floor Β· BELOW_BREAK_EVEN"] - R1 -- no --> R2{losing money after ads?} - R2 -- yes --> UP2["↑ raise Β· LOSING_MONEY"] - R2 -- no --> R3{cover < 35d?} - R3 -- yes --> UP3["↑ +5% Β· LOW_STOCK"] - R3 -- no --> R4{βˆ’30% sales, no cause?} - R4 -- yes --> INV2["πŸ” INVESTIGATE Β· UNEXPLAINED_DROP"] - R4 -- no --> R5{cover > 90d?} - R5 -- yes --> DN["↓ βˆ’5% Β· EXCESS_STOCK"] - R5 -- no --> RC{lost Buy Box AND rival β‰₯3% below?} - RC -- yes --> DNC["↓ toward rival Β· COMPETITOR_UNDERCUT"] - RC -- no --> R6{profit-optimal β‰  current?} - R6 -- yes --> MOVE["↑/↓ toward optimal Β· PROFIT_OPTIMAL"] - R6 -- no --> HOLD["β†’ MAINTAIN Β· NO_SIGNALS"] + COSMOS["COSMOS
sales history Β· costs
stock Β· ad spend"] + RIVALS["Competitor prices
checked on Amazon
not in COSMOS"] + FACTS["1 Β· Work out the facts
what a sale really earns Β·
how long stock lasts Β·
best price we've run"] + CHECK{"2 Β· Run the checklist
stop at the first match"} + FIX["Go fix something
missing costs Β· hidden listing
no price helps yet"] + PRICE["A suggested price
up Β· down Β· leave alone"] + RAILS["3 Β· Safety rails
never below break-even
5% max per step
25% max above today"] + SCREEN["4 Β· Show a person
the price Β· the one reason
the workings Β· our accuracy"] + HUMAN{"5 Β· A human decides"} + OK["βœ… Approve"] + MOD["✏️ Modify"] + NO["βœ–οΈ Reject"] + STOP["Nothing is sent to Amazon
a person makes every change by hand"] + + COSMOS --> FACTS + RIVALS --> FACTS + FACTS --> CHECK + CHECK -->|"pricing can help"| PRICE + CHECK -->|"something is broken"| FIX + PRICE --> RAILS + RAILS --> SCREEN + FIX --> SCREEN + SCREEN --> HUMAN + HUMAN --> OK & MOD & NO + OK & MOD & NO --> STOP + + style COSMOS fill:#d9eee9,stroke:#0c8276,color:#1b2030 + style RIVALS fill:#d9eee9,stroke:#0c8276,color:#1b2030 + style CHECK fill:#fbf6ee,stroke:#a89f8a,color:#1b2030 + style HUMAN fill:#fbf6ee,stroke:#a89f8a,color:#1b2030 + style RAILS fill:#fbe6df,stroke:#df4f33,color:#1b2030 + style FIX fill:#fbe6df,stroke:#df4f33,color:#1b2030 + style STOP fill:#f4f0e8,stroke:#22304e,color:#1b2030 ``` -Guardrails: floor = break-even Γ— 1.05, ceiling = current Γ— 1.25; the recommended move is -capped at Β±5% (bigger steps need elevated approval); the kill switch pauses all approvals. +Two things worth noticing in that picture: -### Competitor rules β€” the two that can move a price, and what bounds them +- **Step 2 can decide that pricing is the wrong tool entirely.** If the cost figures are + missing, or Amazon has hidden our listing, no price change helps β€” so it says so + instead of inventing a number. +- **The safety rails sit between the suggestion and the screen.** Whatever the + calculations produce, nothing that would lose money reaches a person as a + recommendation. -Both branches read a single `CompetitiveState` -([competitive_state.py](src/pricing_agent/competitive_state.py)), never a raw scrape: - -| Rule | Fires when | Effect | -|---|---|---| -| `BUYBOX_SUPPRESSED` | Amazon is not showing our offer | **Investigate, hold.** Placed directly under `NO_COST_DATA`: those are the only two states where the answer is "go and fix something" rather than "set a price". A suppressed variant sells nothing at any price, so its margin and its modelled optimum both describe a listing nobody can buy from. | -| `COMPETITOR_UNDERCUT` | `LOST_PRICE` **and** cheapest rival β‰₯ `competitor_undercut_material_pct` below us | **Decrease toward the rival**, floored and step-capped like every other branch. | - -A third signal, a **competitor premium** while we hold the Buy Box, is a narrative note only. -It never sets a price and never changes an action β€” the elasticity fit is the thing with -evidence behind it, and a premium is not grounds to overrule it. - -Ordering is deliberate: **inventory risk still outranks competitor position, which outranks -profit-optimal.** Chasing a rival down while the shelf is emptying pays margin to sell out -faster. This is visible in the backtest below β€” two of five real SKUs did not move under a -counterfactual undercut precisely because `LOW_STOCK` and `LOSING_MONEY` fired first. - -Three properties make this safe to ship: - -1. **Fail-safe.** Absent, failed, stale (> `competitor_state_max_age_hours`) and - "ownership unknown" all collapse to one flag, and the cascade then computes exactly the - verdict it computed before competitor data existed. Nothing waits on a scrape; nothing is - blocked by one. Competitor data can only ever *add* a verdict. -2. **Never below break-even.** The rival price is a *candidate* (`comp_match`), not a - decision. Verified against real SKUs with a counterfactual rival 60% below us: against a - $14.88 floor the shipped recommendations were $15.19–$21.84, because the Β±5% step cap - binds first. Zero violations. -3. **One named reason per verdict.** No blended scores β€” every fired rule is traceable to a - single reason code, and `logger.info` names the SKU, the rule, the state and the source. - -Thresholds live in [config/pricing_rules.yaml](config/pricing_rules.yaml) -(`competitor_undercut_material_pct: 0.03`, `competitor_premium_material_pct: 0.10`, -`competitor_state_max_age_hours: 6.0`), not in code. The 3% floor sits above the ~2% band our -own realized price already swings through as coupons toggle. - -### Inventory cover matches COSMOS Inventory Planning - -`cover_days` **is COSMOS's own `coverDays`**, so the dashboard and the INVP grid never quote -two different numbers for one SKU. COSMOS counts **inbound** stock against a **7-day** -velocity, so it reads longer than what is on the shelf β€” `UBMICROFIBERDUVETTWINWHITE` is -78 days on (3,999 on hand + 1,030 inbound) Γ· 64/day, against 63 on-hand-only. Both are -reported: the tile leads with the matched figure and appends `63 d on hand, rest inbound`. - -The on-hand figure remains the **fallback**, because COSMOS returns `coverDays: 0` on some -very low-velocity SKUs that hold months of stock (`UBMICROFIBERBS4PCFULLGREY`: 167 units, -334 real days, COSMOS said `0`). Zero satisfies neither inventory rule, so taken literally it -silences both. - -**Trade-off, accepted deliberately:** stockout risk is now judged partly on stock that has not -landed. Measured over the 56-SKU covered line, matching COSMOS moved 7 verdicts β€” -`LOW_STOCK` 7 β†’ 4, `EXCESS_STOCK` 18 β†’ 22. The one to watch is -`UBMICROFIBERBS4PCKINGWHITE`: **12 days on the shelf, 84 with inbound**, so it no longer -raises. If that shipment slips, nothing protects it. - -Display bands are COSMOS's Alpha/Beta scheme (`theme.COVER_BANDS`, Alpha 20/40/70/100). The -pricing **triggers** are separate and live in `pricing_rules.yaml` -(`low_cover_days: 35`, `high_cover_days: 90`) β€” COSMOS's pink at 70 days is a *replenishment* -warning, while crossing a trigger here spends margin on a 5% move. Adopting the band edges -(40/70) would put six more SKUs on a discount; the measured table is in the config beside the -values. - -### Coverage: the comparison sheet gates competitor data, one product line at a time - -The competitor workbook currently covers **one product line**, so the engine reads it as the -first competitor source and **gates on coverage**: - -| SKU | Competitor state | -|---|---| -| In the sheet | Priced from the sheet β€” real like-for-like rival prices, `basis=like-for-like-sheet` | -| Not in the sheet | **`N/A`**, naming what the sheet *does* cover. No rule fires; the verdict is byte-identical to the competitor-blind one | - -**Coverage is the exact SKU set in the sheet, not a line prefix.** Measured against the real -workbook, a prefix gate would be wrong in both directions: the `UBMICROFIBERDUVET` run contains -49 `UBMICROFIBERDUVET*` SKUs **and 7 `UBMICROFIBERBS4PC*`** ones (variants come off the Amazon -parent twister, and COSMOS maps those ASINs to whatever SKU codes they carry), while the line -has 139 SKUs in COSMOS of which only 56 reached a comparison row. So the sheet's own SKU list is -the authority, and "not in the sheet" is reported as a **coverage hole**, never as a claim that -the SKU has no competitors. - -`competitor_sheet_only: true` (default while one line is under test) means an uncovered SKU gets -N/A rather than falling through to a per-ASIN Apify scrape β€” so every verdict either rests on -the sheet or says it has no competitor data. Set it `False` once coverage is broad enough for -Apify to be a sensible fallback. Config: `competitor_sheet_path` (blank = auto-discover the -newest `Competitor_Price_Comparison_*.xlsx`), `competitor_sheet_dirs`, -`competitor_sheet_max_age_hours: 168` (the sheet is a 25–35 min batch run, not a live feed, so -it gets a longer limit than the 6 h single-ASIN one). - -**Two bases, never conflated.** `competitor_min` can be a price of two different things, and the -`basis` field records which: - -- `same-asin-buybox` (Apify) β€” another seller's offer on **our own listing**. Only `LOST_PRICE` - fires the undercut rule; a rival holding the Buy Box *above* us is `LOST_ELIGIBILITY`, where - cutting donates margin. -- `like-for-like-sheet` β€” a **rival brand's** equivalent variant, matched on size + colour. A - cheaper one fires the rule **regardless of who owns our Buy Box**: we can hold ours perfectly - well while a different product undercuts us. This is what the comparison tool exists to - report. It is never labelled a Buy Box loss. - -Two sheet-driven refinements, both from real rows: - -- **A rival whose own Buy Box is suppressed is excluded from the band.** Their price is not - buyable, so undercutting it donates margin for nothing. -- **A material undercut now explains a velocity drop.** Previously a drop with no *own* price - change and no ad collapse was filed `UNEXPLAINED_DROP` even when the sheet held the - explanation β€” observed on `UBMICROFIBERDUVETKINGPURPLE` (rival 15.6% below) and - `UBMICROFIBERBS4PCFULLGREY` (35.2% below). A named cause now converts the Investigate into an - actionable verdict, exactly as the existing stockout branch already did. A stockout still - explains a drop first; an immaterial rival explains nothing. - -Finally, a **suppressed listing has no selling price** (it sells nothing, so COSMOS records no -sales), which used to fail with a bare "no current selling price found in COSMOS". That error now -names the cause, so the most actionable rows in the sheet stop looking like a data problem. - -### Two scrapers, one state - -| Source | Authoritative for | Why | -|---|---|---| -| **Apify** (`tools/amazon/apify.py`) | **Buy Box state read by the engine** | The only source carrying a seller id, so the only one that can tell `WON` from `LOST_PRICE` from `LOST_ELIGIBILITY`. Those lead to opposite actions (cut price vs. fix fulfilment eligibility). | -| **Playwright** (`../scraper/`) | The workbook: like-for-like size/colour matching, BSR, demand buckets, SKU gaps | Apify cannot produce any of it. Its Buy Box field knows only whether a price *rendered*, not whose it was. | - -The split is by **question**, not preference, so neither source is redundant. `reconcile()` -cross-checks the authoritative state against the Playwright run's own cache -(`scraper/.scrape_cache.json`, keyed `ASIN@ZIP`) and **logs any disagreement** rather than -letting a workbook and a recommendation contradict each other in front of a stakeholder. A -disagreement never changes the verdict. Where the authoritative source has nothing usable but -the other has a `SUPPRESSED`, that fact is promoted and the provenance recorded β€” discarding -it to preserve a hierarchy would be choosing the hierarchy over the fact. - -Backtest (`scripts/backtest_competitor_rules.py`, 5 real SKUs, real COSMOS data): - -| Arm | Verdicts changed | -|---|---| -| Real competitor state (both cached entries stale: 66h / 146h) | **0 / 5** β€” the fail-safe working | -| Counterfactual 8% undercut | 3 / 5 β€” the other 2 blocked by `LOW_STOCK` / `LOSING_MONEY` | -| Counterfactual suppression | 5 / 5 β†’ Investigate | -| Prices shipped below break-even, any arm | **0** | - ---- - -## 7. Key formulas - -| Quantity | Formula | -|---|---| -| Take-home / unit | `pΒ·(1 βˆ’ referral% βˆ’ returns%) βˆ’ landed βˆ’ FBA βˆ’ other` | -| Break-even | `(landed + FBA + returns + other) / (1 βˆ’ referral%)` | -| Elasticity | OLS on `ln(units/day) = a + eΒ·ln(price)` over 6 months | -| Scenario demand | `units Γ— (p / pβ‚€)^e` | -| Realization factor | `actual booked profit (window) Γ· modeled net at current price` | -| TACoS | `ad spend Γ· total revenue` (window) | -| Avg sold price | `revenue Γ· units` (window) β€” reconciles the Current row | - ---- - -## 8. Repository map - -``` -pricing_agent/ -β”œβ”€β”€ app.py # dashboard (presentation only) -β”œβ”€β”€ legacy_app.py # previous analyst UI (still runnable) -β”œβ”€β”€ dashboard/ -β”‚ β”œβ”€β”€ theme.py # CRAI design tokens + plotly template -β”‚ └── live_data.py # COSMOS adapter, decision + scenario engine -β”œβ”€β”€ src/pricing_agent/ -β”‚ β”œβ”€β”€ analyze.py # per-SKU orchestration β†’ AnalysisResult -β”‚ β”œβ”€β”€ competitive_state.py # canonical Buy Box state + two-scraper reconciliation -β”‚ β”œβ”€β”€ elasticity.py # demand model + profit optimizer -β”‚ β”œβ”€β”€ performance.py # actual-profit evidence -β”‚ β”œβ”€β”€ tools/margin_engine.py # pure fee/break-even math (golden-tested) -β”‚ └── cosmos/ -β”‚ β”œβ”€β”€ client.py # auth + retry HTTP -β”‚ β”œβ”€β”€ service.py # endpoints, _flatten_takehome, INVP projections -β”‚ └── models.py # typed COSMOS responses (pydantic) -β”œβ”€β”€ config/ # settings + pricing_rules.yaml (incl. competitor thresholds) -β”œβ”€β”€ scripts/ -β”‚ └── backtest_competitor_rules.py # verdict delta, competitor rules blind vs live -β”œβ”€β”€ .streamlit/config.toml # CRAI theme -└── tests/ # incl. margin-engine golden values and - # test_competitor_rules.py (fail-safe + ordering invariants) -``` - ---- - -## 9. Principles - -1. **Deterministic core, narrative shell** β€” every number is a formula over COSMOS data; language models only phrase explanations. -2. **Read-only** β€” the agent proposes; a human approves; nothing writes back. - `submit_price_approval` remains a stub with no callers. -3. **Honest gaps** β€” missing upstream data shows "β€”" or an explicit investigation, never a fabricated number. -4. **Facts vs projections are labeled** β€” the Current row is real booked history; other prices are clearly modeled. -5. **Everything reconciles** β€” one averaging window drives units, revenue, ads and profit so `price Γ— units = revenue` always holds. +Everything on the screen traces back to a specific number from COSMOS. Nothing is +invented along the way. diff --git a/README.md b/README.md index 09ef639..46bf0dc 100644 --- a/README.md +++ b/README.md @@ -37,8 +37,12 @@ when `APIFY_TOKEN` is set: SKU β†’ COSMOS ASIN β†’ scrape `amazon.com/dp/{ASIN}` token the competitive gate stays `UNKNOWN` (COSMOS has no Buy Box data). For a covered product line the **comparison workbook** takes precedence over the per-ASIN -scrape β€” see *Coverage* in [ARCHITECTURE.md](ARCHITECTURE.md), which also documents which of -the two scrapers is authoritative for Buy Box state and why. +scrape β€” see *Coverage* in [ARCHITECTURE-DETAIL.md](ARCHITECTURE-DETAIL.md), which also +documents which of the two scrapers is authoritative for Buy Box state and why. + +New to the project? [ARCHITECTURE.md](ARCHITECTURE.md) is a short, plain-English +explanation of what the tool does and how it decides β€” no coding or Amazon knowledge +assumed. [ARCHITECTURE-DETAIL.md](ARCHITECTURE-DETAIL.md) is the engineering reference. ## Setup diff --git a/app.py b/app.py index df732a1..56e5045 100644 --- a/app.py +++ b/app.py @@ -12,6 +12,7 @@ Run: streamlit run app.py from __future__ import annotations import hmac +import html import os import sys from pathlib import Path @@ -114,9 +115,20 @@ h1, h2, h3, h4 { letter-spacing:-.01em; } [data-testid="stExpander"] summary span.green { color:#0a6f65 !important; } [data-testid="stExpander"] summary [data-testid="stMarkdownContainer"] span[style*="color: rgb(255, 43, 43)"], [data-testid="stExpander"] summary span.red { color:#b23a22 !important; } -/* Stat tiles live in a real auto-fit grid, so they reflow 5 β†’ 4 β†’ 3 β†’ 2 β†’ 1 - across the window width instead of being squeezed into unreadable slivers. */ -.tiles { display:grid; gap:14px; grid-template-columns:repeat(auto-fit, minmax(11.5rem, 1fr)); } +/* Stat tiles reflow with the window instead of being squeezed into slivers. + The column count is keyed to HOW MANY tiles there are, because plain auto-fit + orphans the last one: six tiles in a 980px content column packed 5 + 1, leaving a + lone card against four empty slots, and every note inside the five wrapped to four + lines. Three across gives 3+3 (single SKU) and 3+2 (portfolio) β€” no orphan row and + roughly double the width per note. Only on a genuinely wide window do they open out + to one row, where they still have room to breathe. */ +.tiles { display:grid; gap:14px; grid-template-columns:repeat(3, minmax(0, 1fr)); } +@media (min-width: 1650px) { + .tiles.t5 { grid-template-columns:repeat(5, minmax(0, 1fr)); } + .tiles.t6 { grid-template-columns:repeat(6, minmax(0, 1fr)); } +} +@media (max-width: 1100px) { .tiles { grid-template-columns:repeat(2, minmax(0, 1fr)); } } +@media (max-width: 560px) { .tiles { grid-template-columns:1fr; } } .tl { background:#fffdf9; border:1px solid #e3ddd0; border-radius:14px; padding:16px 18px; box-shadow:0 1px 2px rgba(27,32,48,.05), 0 8px 24px rgba(27,32,48,.06); } .tl-ic { width:36px; height:36px; border-radius:10px; background:#d9eee9; display:flex; @@ -143,13 +155,26 @@ h1, h2, h3, h4 { letter-spacing:-.01em; } .brandlogo .t2 { font-size:.72rem; color:#7c8092; } .invt-wrap { overflow-x:auto; border:1px solid #e3ddd0; border-radius:12px; background:#fffdf9; } .invt { border-collapse:collapse; width:100%; font-size:.78rem; } +/* Numeric columns are RIGHT-aligned with tabular (fixed-width) figures, so digits + line up by place value down the column and magnitudes can be compared at a glance. + Centred money was the default here, which staggers the decimal point on every row + and makes $9,801 and $981 look the same length. Headers follow their column. + `.tnum` opts a cell out (the first column, and any text cell). */ +.invt { font-variant-numeric: tabular-nums; font-feature-settings:"tnum" 1, "lnum" 1; } .invt th { background:#faf7f1; padding:9px 10px; color:#4a4f60; font-weight:600; - border-bottom:1px solid #e3ddd0; white-space:nowrap; text-align:center; } + border-bottom:1px solid #e3ddd0; white-space:nowrap; text-align:right; } .invt th:first-child, .invt td:first-child { text-align:left; position:sticky; left:0; background:#fffdf9; box-shadow:1px 0 0 #e3ddd0; min-width:215px; } .invt th:first-child { background:#faf7f1; } -.invt td { padding:9px 10px; border-left:1px solid #ece6d9; text-align:center; +.invt td { padding:9px 10px; border-left:1px solid #ece6d9; text-align:right; min-width:98px; vertical-align:middle; } +/* Text cells (labels, prose advice) stay left-aligned β€” right-aligning a sentence + gives it a ragged left edge, which is much harder to read than ragged right. */ +.invt th.txt, .invt td.txt { text-align:left; } +/* Sub-values stacked under a figure (the "avg sold" / "βœ“ ran Nd" provenance lines, + and the inventory cell's value line) inherit the column's right edge, so the + number and its caption share one alignment spine. */ +.invt .cell-val, .invt .p-s { text-align:inherit; } .invt .p-t { font-weight:700; color:#1b2030; font-size:.82rem; } .invt .p-s { color:#7c8092; font-size:.71rem; margin-top:2px; white-space:nowrap; } .invt .cell-main { font-weight:800; font-size:.93rem; color:#1b2030; } @@ -226,7 +251,10 @@ h1, h2, h3, h4 { letter-spacing:-.01em; } min-width:42% !important; } [class*="st-key-pillrow"] [data-testid="stColumn"] { min-width:8rem !important; } - .tiles { grid-template-columns:repeat(auto-fit, minmax(9.5rem, 1fr)); gap:10px; } + /* Two across, not auto-fit: in a ~485px column auto-fit still packed THREE tiles + at ~150px each, which is narrower than the same tiles get on a desktop and left + every note wrapping to four lines. Two is the widest count that stays legible. */ + .tiles { grid-template-columns:repeat(2, minmax(0, 1fr)) !important; gap:10px; } .loadcard { margin-top:6vh; padding:22px 18px; } .loadcard .lc-b { font-size:.72rem; letter-spacing:-.5px; } } @@ -347,10 +375,23 @@ def _parse_skus(raw: str) -> tuple: # ---------------------------------------------------------------- helpers def esc(s: str) -> str: - """Escape $ so Streamlit markdown never enters LaTeX math mode.""" + """Escape $ so Streamlit markdown never enters LaTeX math mode. + + MARKDOWN ONLY. Inside a raw HTML block this is the wrong tool β€” Streamlit does not + run LaTeX there, so it would render a literal backslash. Use `esc_html` instead. + """ return s.replace("$", "\\$") +def esc_html(s) -> str: + """Escape a value for interpolation into one of the raw HTML tables. + + Amazon listing titles routinely carry `&`, quotes and the odd angle bracket; left + raw they either swallow the rest of the cell or surface as a stray entity. + """ + return html.escape(str(s), quote=True) + + def dash(v, fmt: str = "{}") -> str: """Format a value, or an em-dash when it is missing.""" if v is None or (isinstance(v, float) and pd.isna(v)): @@ -358,11 +399,34 @@ def dash(v, fmt: str = "{}") -> str: return fmt.format(v) +# U+2212 MINUS SIGN, not the ASCII hyphen. It is the same width as the digits in a +# tabular-numeral column so negative rows stay aligned, and it matches the "βˆ’2%" / +# "βˆ’5%" scenario labels the engine already emits. Used for every negative money and +# percentage in the UI so one screen never shows three different minus glyphs. +MINUS = "βˆ’" + + +def money(v: float, dp: int = 0) -> str: + """Signed currency with the minus OUTSIDE the symbol: βˆ’$108, not $-108. + + Python's `f"${v:,.0f}"` puts the sign where the digits start, which reads as a + dollar sign applied to a negative number rather than a negative amount, and in a + right-aligned column it also shunts the $ out of line with the rows above it. + """ + return f"{MINUS if v < 0 else ''}${abs(v):,.{dp}f}" + + +def pct(v: float, dp: int = 1, signed: bool = True) -> str: + """Percentage with a real minus sign and an explicit + when signed.""" + sign = MINUS if v < 0 else ("+" if signed else "") + return f"{sign}{abs(v):.{dp}f}%" + + def compact_money(v: float) -> str: """$33k rather than $32,974 β€” false precision from a model with a known error band reads as certainty the number does not have.""" a = abs(v) - sign = "-" if v < 0 else "" + sign = MINUS if v < 0 else "" if a >= 1_000_000: return f"{sign}${a / 1_000_000:.1f}M" if a >= 1_000: @@ -386,6 +450,27 @@ def money_range(v: float, err: float | None) -> str: return f"{compact_money(min(lo, hi))}–{compact_money(max(lo, hi))}" +def econ_rows(d: dict) -> tuple[dict, dict]: + """(economics at TODAY's price, economics at the price we recommend TODAY). + + The second one is deliberately matched on `rec_price`, NOT on `rec_key`. When a + move is step-capped the two are different prices: `rec_key` names the rung the + cascade chose β€” the $20.00 destination β€” while the headline recommends $17.94 on + the way there. Quoting the destination's units, profit and margin beside a $17.94 + headline promises what that move does not deliver. + + Both the header tiles and the queue decision card read this, so they cannot drift: + the tiles were fixed for this and the card was not, and the card sat directly under + "↑ RAISE $17.09 β†’ $17.94" showing the economics of $20.00. + """ + econ = d.get("scen_econ") or [] + by_key = {x["key"]: x for x in econ} + cur = by_key.get("current") or {} + rec = next((x for x in econ if abs(x["price"] - d["rec_price"]) < 0.011), + by_key.get(d.get("rec_key")) or {}) + return cur, rec + + def unit_take_home(price: float, d: dict) -> float: """Per-unit take-home under this SKU's COSMOS fee model.""" rp = d.get("referral_pct", REFERRAL_PCT) @@ -398,9 +483,9 @@ def _price_move_html(delta_pct: float) -> str: """Colored arrow + percentage chip for a price move vs current.""" if abs(delta_pct) < 0.05: return 'β†’ 0.0%' - if delta_pct > 0: - return (f'β–² +{delta_pct:.1f}%') - return (f'β–Ό {delta_pct:.1f}%') + colour = "#0a6f65" if delta_pct > 0 else "#c9442b" + return (f'' + f'{"β–²" if delta_pct > 0 else "β–Ό"} {pct(delta_pct)}') WINDOW_OPTS = {"7 days": 7, "14 days": 14, "30 days": 30, "90 days": 90, @@ -491,8 +576,8 @@ def render_scenarios(d: dict, key_prefix: str, compact: bool = False): a1, a2, a3 = st.columns(3) a1.metric(f"Units/day (last {days}d)", f"{cur['units_day']:,}", f"avg sold ${price_check:,.2f}/unit", delta_color="off") - a2.metric("Net profit /30d", f"${cur['net_30d']:,.0f}", - f"on ${cur['revenue_30d']:,.0f} revenue", delta_color="off") + a2.metric("Net profit /30d", money(cur["net_30d"]), + f"on {money(cur['revenue_30d'])} revenue", delta_color="off") if best: lbl = "Best price = hold" if best["key"] == "current" else "Best price (projected)" a3.metric(lbl, f"${best['price']:.2f}", @@ -528,18 +613,18 @@ def render_scenarios(d: dict, key_prefix: str, compact: bool = False): move_cell = _price_move_html(x["delta_pct"]) rows_html.append( f'' - f'{emoji} {esc(x["label"])}{star}' + f'{emoji} {esc_html(x["label"])}{star}' f'{price_cell}' f'{move_cell}' f'{x["units_day"]:,}' - f'${x["revenue_30d"]:,.0f}' - f'${x["ad_30d"]:,.0f}' - f'${x["net_30d"]:,.0f}' + f'{money(x["revenue_30d"])}' + f'{money(x["ad_30d"])}' + f'{money(x["net_30d"])}' f'{x["net_margin_pct"]:.1f}%' # No esc() inside the raw HTML table β€” Streamlit does not run LaTeX on # these cells, so escaping would render a literal backslash. + ("" if compact else - f'{x["advice"]}') + f'{x["advice"]}') + '' ) # Column headers with hover tooltips explaining how each number is computed. @@ -553,8 +638,10 @@ def render_scenarios(d: dict, key_prefix: str, compact: bool = False): tacos_now = cur["ad_30d"] / cur["revenue_30d"] * 100 if cur["revenue_30d"] else 0.0 def th(label, tip, left=False): - align = "left" if left else "center" - return (f'' + """Header cell. Numeric headers inherit the stylesheet's right alignment so + they sit over their column's digits; text headers opt out via `.txt`.""" + cls = ' class="txt"' if left else "" + return (f'' f'{label}' f' β“˜') @@ -951,8 +1038,8 @@ def inventory_table_html(d) -> str: '
' f'{heads}' "" - f'' @@ -1152,7 +1239,7 @@ def _on_progress(frac: float, message: str): history, elasticity & bulk economics
{bar}
{pct}%
-
{esc(message)}
+
{esc_html(message)}
""", unsafe_allow_html=True, ) @@ -1294,15 +1381,9 @@ if len(details) == 1: _sku = next(iter(details)) _d = details[_sku] _e1 = model_error_pct(_d) - _econ = {x["key"]: x for x in (_d.get("scen_econ") or [])} - _cur_e = _econ.get("current") or {} - # The row at the price we are RECOMMENDING TODAY, not the rung the cascade named. When a - # move is step-capped those differ: `rec_key` here is the $20.00 destination while the - # headline recommends $17.94, and quoting the destination's margin beside a $17.94 headline - # would promise 13.4% for a move that delivers 6.6%. - _rec_e = next((x for x in (_d.get("scen_econ") or []) - if abs(x["price"] - _d["rec_price"]) < 0.011), - _econ.get(_d.get("rec_key")) or {}) + # Anchored on the price we RECOMMEND TODAY, not the step-capped destination the + # cascade named β€” see econ_rows, which the queue decision card shares. + _cur_e, _rec_e = econ_rows(_d) _fi = _d.get("floor_info") or {} _floor = _fi.get("floor") _cover = _d.get("cover_days") @@ -1393,26 +1474,59 @@ if len(details) == 1: tiles = [_t_impact, _t_margin, _t_floor, _t_cover, _t_vel, _t_comp] else: + # These five tiles are NOT all counted over the same set, and reading them as one + # row is how "4 SKUs needing action" ends up sitting beside "12 overstock-risk" as + # if 8 had been dismissed. The first three are the PENDING queue β€” work still to do. + # The last two are catalogue facts that hold whether or not a recommendation has + # been actioned. Every note names its own denominator so the row cannot be misread. + _n_all = len(details) tiles = [ - ("⚑", str(len(needing_action)), "SKUs needing action", "queue below, ranked by impact"), + ("⚑", str(len(needing_action)), "SKUs needing action", + f"of {len(pending)} pending Β· queue below, ranked by impact"), ("πŸ’°", money_range(opportunity, _err), "Profit opportunity Β· 30d", - (f"open positive impacts Β· Β±{_err * 100:.0f}% model error" if _err - else "sum of open positive impacts")), - ("⏳", str(len(pending)), "Pending approvals", f"{len(bulk_eligible)} bulk-eligible"), - ("πŸ“‰", str(stockout_risk), "Stockout-risk SKUs", f"cover ≀ {LOW_COVER_DAYS} days"), - ("πŸ“¦", str(overstock_risk), "Overstock-risk SKUs", f"cover β‰₯ {HIGH_COVER_DAYS} days"), + (f"across {len(pending)} pending Β· Β±{_err * 100:.0f}% model error" if _err + else f"sum of positive impacts across {len(pending)} pending")), + ("⏳", str(len(pending)), "Pending approvals", + f"of {_n_all} loaded Β· {len(bulk_eligible)} bulk-eligible"), + ("πŸ“‰", str(stockout_risk), "Stockout-risk SKUs", + f"of {_n_all} loaded Β· cover ≀ {LOW_COVER_DAYS} days"), + ("πŸ“¦", str(overstock_risk), "Overstock-risk SKUs", + f"of {_n_all} loaded Β· cover β‰₯ {HIGH_COVER_DAYS} days"), ] -st.markdown('
' +st.markdown(f'
' # Tiles are (icon, value, label, note) with an OPTIONAL 5th accent colour, so a # banded reading can tint its icon without every other tile growing a field. + # The `t{n}` class lets the grid pick a column count that divides evenly and + # never strands the last tile on a row of its own. + "".join(theme.tile(*t) for t in tiles) + '
', unsafe_allow_html=True) st.markdown("") +# ---------------------------------------------------------------- queue filters +# Applied BEFORE the pills are drawn, because the pill counts have to describe the +# list the pills actually produce. They used to be counted on the whole catalogue +# while the queue below was filtered by marketplace, trust, status and search β€” so +# with a marketplace selected "↑ Raise Β· 12" opened a list of three. +# +# The action filter is deliberately NOT part of this base: a facet count must not be +# narrowed by the facet it is counting, or every unselected pill reads 0. +base_rows = summary.copy() +if f_mkt != "All": + base_rows = base_rows[base_rows["marketplace"] == f_mkt] +if f_conf != "All": + base_rows = base_rows[base_rows["trust_tier"] == f_conf] +if f_stat != "All": + base_rows = base_rows[[st.session_state.status[s] == f_stat for s in base_rows["sku"]]] +if q: + base_rows = base_rows[ + base_rows["sku"].str.contains(q, case=False, regex=False) + | base_rows["title"].str.contains(q, case=False, regex=False)] + # ---------------------------------------------------------------- action pills -counts = summary["action"].value_counts().to_dict() -pills = [("All", "All", len(summary)), +counts = base_rows["action"].value_counts().to_dict() +_filtered = len(base_rows) != len(summary) +pills = [("All", "All", len(base_rows)), ("Increase", "↑ Raise", counts.get("Increase", 0)), ("Decrease", "↓ Lower", counts.get("Decrease", 0)), ("Maintain", "β†’ Hold", counts.get("Maintain", 0)), @@ -1420,25 +1534,21 @@ pills = [("All", "All", len(summary)), with st.container(key="pillrow"): # keyed so the CSS can keep these pills tight pcols = st.columns([0.8, 1, 1, 1, 1, 2]) for col, (a, lbl, n) in zip(pcols, pills): + _scope = " matching the current filters" if _filtered else "" col.button(f"{lbl} Β· {n}", key=f"pill_{a}", type="primary" if st.session_state.action_sel == a else "secondary", use_container_width=True, on_click=set_action, args=(a,), - help=f"Show {a.lower()} recommendations" if a != "All" else "Show all") + help=(f"Show all {len(base_rows)} recommendations{_scope}" if a == "All" + else f"Show the {n} {a.lower()} recommendation" + f"{'s' if n != 1 else ''}{_scope}")) +if _filtered: + st.caption(f"Counts reflect the {len(base_rows)} of {len(summary)} SKUs matching the " + f"sidebar filters.") # ---------------------------------------------------------------- queue -rows = summary.copy() -if f_mkt != "All": - rows = rows[rows["marketplace"] == f_mkt] -if f_conf != "All": - rows = rows[rows["trust_tier"] == f_conf] -if f_stat != "All": - rows = rows[[st.session_state.status[s] == f_stat for s in rows["sku"]]] +rows = base_rows if st.session_state.action_sel != "All": rows = rows[rows["action"] == st.session_state.action_sel] -if q: - mask = (rows["sku"].str.contains(q, case=False, regex=False) - | rows["title"].str.contains(q, case=False, regex=False)) - rows = rows[mask] if rows.empty: st.info("No recommendations match the current filters.") @@ -1542,13 +1652,13 @@ for _, r in rows.iterrows(): f"pricing to it loses money on every unit.")) with c2: # Calibrated net profit for the headline metric (matches the tables). - _econ = {x["key"]: x for x in (d.get("scen_econ") or [])} - _rec_e = _econ.get(d["rec_key"]) - _cur_e = _econ.get("current") - rec_units = _rec_e["units_day"] if _rec_e else round(d["rec_units_day"]) - cur_units = _cur_e["units_day"] if _cur_e else round(d["units_day"]) - rec_net = _rec_e["net_30d"] if _rec_e else d["rec_profit_30d"] - cur_net = _cur_e["net_30d"] if _cur_e else d["profit_30d"] + # Anchored on the price this card RECOMMENDS, not the step-capped + # destination β€” see econ_rows. + _cur_e, _rec_e = econ_rows(d) + rec_units = _rec_e.get("units_day", round(d["rec_units_day"])) + cur_units = _cur_e.get("units_day", round(d["units_day"])) + rec_net = _rec_e.get("net_30d", d["rec_profit_30d"]) + cur_net = _cur_e.get("net_30d", d["profit_30d"]) m1, m2, m3, m4 = st.columns(4) m1.metric("Units/day", f"{rec_units:,}", f"{rec_units - cur_units:+,} vs now", delta_color="off") @@ -1708,7 +1818,11 @@ for _, r in rows.iterrows(): f"**${d['current_price']:,.2f}**, and every option below is " f"calculated from that, so an option can sit below the Current " f"row without being a price cut.") - st.caption(cap) + # esc(): the dollar amounts above would otherwise open LaTeX math + # mode and Streamlit would render the whole sentence as run-together + # italic maths ("18.30 * *β€”theaveragepriceactually*sold*over..."). + # The markdown emphasis in `cap` survives β€” esc only touches "$". + st.caption(esc(cap)) with oc2: st.button("πŸ“Š View more", key=f"vm_{sku}", use_container_width=True, on_click=goto_view, args=(sku, VIEW_SCENARIOS), @@ -1898,7 +2012,9 @@ for _, r in rows.iterrows(): # listing. Calling the latter "an offer on this ASIN" would be simply untrue. if meta.get("source") == "comparison-sheet": as_of = meta.get("sheet_as_of") - st.caption( + # esc(): two dollar amounts on one line make Streamlit treat the text + # between them as LaTeX and render it as run-together italic maths. + st.caption(esc( f"**{len(rivals)} like-for-like rival(s)** from the comparison sheet β€” " f"each is a *different brand's own listing* matched on size + colour, " f"not an offer on your ASIN" @@ -1906,12 +2022,13 @@ for _, r in rows.iterrows(): if meta.get("exact_rivals") is not None else "") + (f" Β· median ${meta['median']:,.2f}" if meta.get("median") else "") + f" Β· your price ${d['current_price']:.2f}" - + (f" Β· sheet dated {as_of:%d %b %Y}" if as_of else "") + ".") + + (f" Β· sheet dated {as_of:%d %b %Y}" if as_of else "") + ".")) else: - st.caption(f"**{len(rivals)} competitor offer(s)** on this ASIN" - + (f" Β· competitor median ${meta['median']:,.2f}" - if meta.get("median") else "") - + f" Β· your price ${d['current_price']:.2f}.") + st.caption(esc( + f"**{len(rivals)} competitor offer(s)** on this ASIN" + + (f" Β· competitor median ${meta['median']:,.2f}" + if meta.get("median") else "") + + f" Β· your price ${d['current_price']:.2f}.")) disp = pd.DataFrame({ "Seller": rivals["seller"], "ASIN": rivals["asin"].map(lambda a: _amazon_url(a, r["marketplace"])), diff --git a/claude.md b/claude.md index 8a3d527..17a9c18 100644 --- a/claude.md +++ b/claude.md @@ -129,7 +129,10 @@ separate "Competitors" display tab. Change this deliberately and narrowly: otherwise. - Any new number that ends up in a scenario/recommendation must be traceable to a named source field, exactly like the existing data dictionary β€” update the docs - (`ARCHITECTURE.md` / the README) with any new field or rule you add. + (`ARCHITECTURE-DETAIL.md` β€” the engineering reference β€” / the README) with any new + field or rule you add. `ARCHITECTURE.md` is the plain-English overview for + non-technical readers; keep it short and jargon-free, and only touch it when the + behaviour a non-specialist would care about actually changes. - Don't add scope beyond this: no new sixth COSMOS endpoint, no new speculative feature, unless it's required to fix a bug or complete the wiring above. diff --git a/config/pricing_rules.yaml b/config/pricing_rules.yaml index fddddb1..275a5e7 100644 --- a/config/pricing_rules.yaml +++ b/config/pricing_rules.yaml @@ -74,6 +74,29 @@ competitor_rules_enabled: true # distinguished from our own price wobble. competitor_undercut_material_pct: 0.03 +# Whether a LIKE-FOR-LIKE SHEET undercut needs a second, independent signal before it may cut +# a price. (Same-ASIN undercuts are unaffected: LOST_PRICE already IS the corroboration β€” +# Amazon has taken the Buy Box off us over price.) +# +# A sheet row says a different brand's comparable product is cheaper. That is a true fact about +# the market and worth reporting, but on its own it is not evidence the rival price is costing +# us anything: we can be 3% dearer, hold our own Buy Box, and sell perfectly well on brand, +# reviews or Prime badge. Cutting there spends certain margin against an unmeasured threat, and +# it fired ABOVE the profit-optimal rule, so it could overrule an elasticity fit that wanted a +# RAISE β€” on a healthy, Buy-Box-winning SKU. +# +# With this on, a sheet undercut cuts only when something else agrees: +# * demand has materially dropped (the 30d vs 6m VELOCITY_DROP test), or +# * we are not actually winning the Buy Box. +# Both of those mean the cheaper rival is a plausible explanation for something we can measure. +# +# The documented fall-through from UNEXPLAINED_DROP is preserved: those SKUs +# (UBMICROFIBERDUVETKINGPURPLE, UBMICROFIBERBS4PCFULLGREY) reach the competitor branch WITH a +# velocity drop, so they still get COMPETITOR_UNDERCUT rather than being filed unexplained. +# +# Set false to restore the previous behaviour (any material sheet undercut cuts). +competitor_sheet_requires_corroboration: true + # How far ABOVE the rival field we have to sit before it is worth remarking on while we # still hold the Buy Box. Only ever a narrative note -- it never moves a price on its own. competitor_premium_material_pct: 0.10 diff --git a/config/settings.py b/config/settings.py index 4e5b6c8..aacf66b 100644 --- a/config/settings.py +++ b/config/settings.py @@ -36,6 +36,7 @@ class PricingRules(BaseModel): # sidebar's global approval pause. competitor_rules_enabled: bool = True competitor_undercut_material_pct: float = 0.03 + competitor_sheet_requires_corroboration: bool = True competitor_premium_material_pct: float = 0.10 competitor_state_max_age_hours: float = 6.0 competitor_sheet_max_age_hours: float | None = 168.0 diff --git a/dashboard/live_data.py b/dashboard/live_data.py index 7b65323..9d5199a 100644 --- a/dashboard/live_data.py +++ b/dashboard/live_data.py @@ -358,6 +358,11 @@ def price_floor(r, cur_price: float) -> dict: MATERIAL_GAIN = 1.10 # observed profit must beat today's by this much to act +# Findings that outrank the AD_SPIRAL verdict. AD_SPIRAL is applied after the cascade and +# overrides whatever fired, so this set is what it must NOT override. See the block in +# `_decide` for why each one earns its place. +AD_SPIRAL_YIELDS_TO = frozenset({"NO_COST_DATA", "BUYBOX_SUPPRESSED", "LOW_STOCK"}) + # A day counts as "effectively out of stock" below this many days of cover. # NOT zero: Amazon's feed never reports a clean 0 β€” the audited SKU had 90/90 days # of inventory readings and not one at zero β€” so an `inventory <= 0` test would @@ -413,8 +418,16 @@ def _evidence_target(r, cur_price: float, keys: set) -> tuple | None: return None # not materially better than today if bday <= 0: return None # every observed price lost money - days = next((b["days"] for b in (r.price_bands or []) - if abs(b["avg_price"] - bop) < 0.01), 0) + # Sample size comes from the analysis, which read it straight off the winning band. + # It used to be re-derived here by matching `bop` against each band's `avg_price` β€” + # but `bop` was the band's $0.50-rounded KEY, which never equals the weighted + # `avg_price`, so this always fell through to 0 and the rationale below read + # "booked $X/day of actual profit over 0 days". The fallback keeps the old lookup + # for callers that predate the field; it now compares like with like. + days = getattr(r, "best_observed_days", None) + if not days: + days = next((b["days"] for b in (r.price_bands or []) + if abs(b["avg_price"] - bop) < 0.01), 0) gain = bday - cur_day if cur_day is not None else bday why = (f"${bop:.2f} booked ${bday:,.0f}/day of actual profit over {days} days" + (f", against ${cur_day:,.0f}/day at today's price" if cur_day is not None @@ -449,10 +462,27 @@ def _decide(r, cur_price: float, hist: pd.DataFrame, fp: dict, scen: pd.DataFram # as price, each extra $1 of price buys only cents of contribution, so no # price reaches break-even β€” recommending a raise would cost volume and fix # nothing. The lever is ad efficiency or COGS. - # Both this and BUYBOX_SUPPRESSED return Investigate-and-hold, so neither moves a price - # either way β€” but if the Buy Box is suppressed that is the thing to go and fix, and it - # must not be relabelled as an advertising problem on the way out. - if r.ad_curve_unrecoverable and "BUYBOX_SUPPRESSED" not in reasons: + # + # This runs AFTER the cascade and overrides whatever fired, so the set below is the list + # of findings that outrank it. Each is here for its own reason: + # + # NO_COST_DATA the ad-curve maths is not even trustworthy without costs β€” `fixed` is + # understated when costPerUnit/fbaFee are 0, which makes `contribution + # <= 0` EASIER to hit. So a missing-COGS SKU could be told its problem + # was advertising, sending someone to the ad console when the fix is a + # data-entry field. Diagnosing off known-bad inputs is worse than not + # diagnosing. + # BUYBOX_SUPPRESSED a listing nobody can buy from has no meaningful ad economics, and + # the thing to go and fix is the suppression. + # LOW_STOCK this one changes a PRICE, not just a label. The low-stock branch raises + # 5% to slow the burn on a shelf that is about to empty, and that is right + # whatever the ad slope does β€” the units are going to sell regardless, so + # the only question is what we get for them. Holding instead sells the + # last of the stock cheap. + # + # Everything else still yields to AD_SPIRAL: where price genuinely cannot reach break-even, + # naming that beats recommending a move that fixes nothing. + if r.ad_curve_unrecoverable and not (set(reasons) & AD_SPIRAL_YIELDS_TO): cpd = r.contribution_per_dollar detail = (f"each extra $1 of price yields only ${cpd:.2f} of contribution" if cpd is not None else "ad cost rises faster than price") @@ -620,9 +650,25 @@ def _decide_raw(r, cur_price: float, hist: pd.DataFrame, fp: dict, scen: pd.Data # hold our own perfectly well while a different product undercuts us. It is exactly # what the comparison workbook exists to report, and the sibling tool's own action # classifier already treats it this way. + # + # A SHEET undercut additionally needs CORROBORATION before it may cut. On its own it + # says only "a different brand is cheaper" β€” true, reportable, but not evidence that + # the rival price is costing us anything. We can sit 3% dearer, hold our own Buy Box + # and sell fine on brand, reviews or the Prime badge; cutting there spends certain + # margin against an unmeasured threat. And because this branch sits ABOVE + # PROFIT_OPTIMAL, it could overrule an elasticity fit that wanted a RAISE. + # + # Corroboration = something else we can measure agrees the rival matters: + # * demand has materially dropped (the same `drop` the VELOCITY_DROP tests use), or + # * we are not actually winning the Buy Box. + # The documented UNEXPLAINED_DROP fall-through is unaffected β€” those SKUs arrive here + # WITH `drop` set, which is exactly the first form of corroboration. + sheet_corroborated = bool(drop or comp.status is not BuyBoxStatus.WON) + sheet_qualifies = comp.basis == BASIS_SHEET and ( + sheet_corroborated or not rules.competitor_sheet_requires_corroboration) basis_qualifies = ( (comp.basis == BASIS_SAME_ASIN and comp.status is BuyBoxStatus.LOST_PRICE) - or comp.basis == BASIS_SHEET + or sheet_qualifies ) comp_undercut = bool( basis_qualifies @@ -630,6 +676,15 @@ def _decide_raw(r, cur_price: float, hist: pd.DataFrame, fp: dict, scen: pd.Data and comp_gap >= rules.competitor_undercut_material_pct and comp_price ) + # Say so when a real, material undercut was seen and deliberately not acted on β€” + # otherwise this looks identical to "no competitor data" in the root cause. + if (comp.basis == BASIS_SHEET and not sheet_qualifies and comp_gap is not None + and comp_gap >= rules.competitor_undercut_material_pct and comp_price): + root.append(( + "Competitor undercut not acted on", + f"a like-for-like rival is {comp_gap:.1%} cheaper (${comp_price:.2f}), but we " + f"hold the Buy Box and demand has not dropped β€” no evidence the gap is costing " + f"us volume, so no price cut. Reported, not acted on.")) root.append(("Competitor position", f"{comp.status.value} ({comp.source}, basis={comp.basis}" + (f", {comp.rivals} rival(s)" if comp.rivals else "") @@ -971,6 +1026,10 @@ SCEN_LABELS = { "current": "Current", "up_2": "+2%", "up_5": "+5%", "down_2": "βˆ’2%", "down_5": "βˆ’5%", "match_median": "Match competitor", "min_safe": "Safe floor", "max_profit": "Profit-optimal", "best_observed": "Best observed", + # Both of these are injected into the grid by build_live_sku after the dict above + # was written, so they had no entry and `SCEN_LABELS.get(key, key)` fell through to + # the raw key β€” the scenarios table showed a row literally labelled "comp_match". + "comp_match": "Match cheapest rival", "recommended": "Recommended", } @@ -1259,6 +1318,28 @@ def reproject_inventory(projections: list, units_day_now: float, units_day_new: return rows +def projection_elasticity(fit: dict | None) -> tuple[float, bool]: + """The slope every PROJECTION is built on, and whether it is this SKU's own. + + Scenario units, the 30-day impact tile, the portfolio opportunity total and the + queue sort all trace back here, so it must honour the SAME ``actionable`` gate the + decision does. It used to take the raw point estimate whenever one existed, which + let two unusable fits through: + + * a slope whose 95% CI spans zero β€” no established relationship at all; + * a POSITIVE slope, which `estimate_elasticity` can return (``actionable`` + requires e < 0, the returned value is not clamped) and which projects that + RAISING price sells MORE units. Every scenario above today's price then reads + as free money, and the queue sorts the worst SKUs to the top. + + Falling back to FALLBACK_ELASTICITY is the same conservative assumption the + decision already makes when it gates `max_profit` out. + """ + f = fit or {} + fitted = bool(f.get("actionable") and f.get("elasticity")) + return (f["elasticity"] if fitted else FALLBACK_ELASTICITY), fitted + + def _observed_band(price: float, bands: list | None, tol: float = 0.02): """The observed price band matching `price`, when this price has really run.""" near = [b for b in (bands or []) @@ -1518,7 +1599,14 @@ def build_live_sku(svc, sku: str, with_competitive: bool = False, comp_meta["sheet_line"] = sheet.line if sheet is not None else None comp_meta["sheet_covered_skus"] = len(sheet.rows) if sheet is not None else 0 - el = (r.elasticity or {}).get("elasticity") or FALLBACK_ELASTICITY + # Gated on `actionable` β€” see projection_elasticity for why an ungated point + # estimate could put a positive slope behind every number on screen. + _el_fit = r.elasticity or {} + el, el_fitted = projection_elasticity(_el_fit) + if not el_fitted and _el_fit.get("elasticity"): + logger.info("%s: elasticity %.2f not usable (%s) β€” projections fall back to the " + "default %.2f", sku, _el_fit["elasticity"], + _el_fit.get("why") or "not actionable", FALLBACK_ELASTICITY) be = r.break_even or 0.0 # The floor that actually matters: the highest of accounting / ad-inclusive / # empirical break-even. Everything below it loses money on every unit. @@ -1804,6 +1892,11 @@ def build_live_sku(svc, sku: str, with_competitive: bool = False, observed_price_min=r.observed_price_min, observed_price_max=r.observed_price_max, elasticity_actionable=r.elasticity_actionable, elasticity_detail=r.elasticity, + # The slope the projections on screen were ACTUALLY built with, and whether it + # is this SKU's own fitted value or the conservative default. Without the flag, + # a scenario table built on FALLBACK_ELASTICITY is indistinguishable from one + # built on a measured response. + elasticity_used=el, elasticity_is_fitted=el_fitted, profit_optimal_blocked_reason=r.profit_optimal_blocked_reason, profit_optimal_unconstrained=r.profit_optimal_unconstrained, ad_cost_model=r.ad_cost_model, price_bands=r.price_bands, @@ -1836,6 +1929,7 @@ def build_live_sku(svc, sku: str, with_competitive: bool = False, unprofitable_months=r.unprofitable_months, best_observed_price=r.best_observed_price, best_observed_profit_day=r.best_observed_profit_day, + best_observed_days=getattr(r, "best_observed_days", None), unmodeled_cost_gap=r.unmodeled_cost_gap, suggested_price=r.suggested_price, trend=r.trend, diff --git a/dashboard/theme.py b/dashboard/theme.py index 76d23d1..6b61628 100644 --- a/dashboard/theme.py +++ b/dashboard/theme.py @@ -4,6 +4,8 @@ Palette lifted from crai.utopiabrands.com (auth.css / app.css): warm cream paper off-white cards, teal primary, coral accent, navy chrome, Inter type. """ +from html import escape + import plotly.graph_objects as go import plotly.io as pio @@ -123,11 +125,16 @@ def tile(icon: str, value: str, label: str, note: str = "", `accent` tints the icon square only β€” a band colour marks the reading without repainting the card, so the tiles still read as one set and the colour never becomes the only signal (the band's word travels in `note`). + + Every caller-supplied string is HTML-escaped: `value` and `note` carry live data + (a Buy Box state, a product-derived label), and an unescaped `&` there renders as a + broken entity mid-tile. """ - note_html = f'
{note}
' if note else "" + value, label = escape(str(value)), escape(str(label)) + note_html = f'
{escape(str(note))}
' if note else "" ic_style = f' style="background:{accent}"' if accent else "" return ( - f'
{icon}
' + f'
{escape(str(icon))}
' f'
{value}
{label}
{note_html}
' ) diff --git a/scripts/backtest_competitor_rules.py b/scripts/backtest_competitor_rules.py index fc6c621..b60597b 100644 --- a/scripts/backtest_competitor_rules.py +++ b/scripts/backtest_competitor_rules.py @@ -35,6 +35,7 @@ from __future__ import annotations import argparse import sys +from contextlib import contextmanager from datetime import datetime, timezone from pathlib import Path @@ -59,19 +60,35 @@ DEFAULT_SKUS = [ def capture_decide(): - """Wrap _decide so we keep the real arguments it was called with.""" + """Wrap _decide so we keep the real arguments it was called with. + + Signature-agnostic on purpose. This spy was pinned to `_decide`'s exact parameter list + and silently rotted when the cascade gained `cover_days` β€” the wrapper then raised + TypeError on the first SKU and the whole backtest was unrunnable. Capturing *args/**kw + and replaying them verbatim means the harness cannot drift from the function under test. + """ calls: list[dict] = [] original = L._decide - def spy(r, cur_price, hist, fp, scen, outlook=None, comp=None): - calls.append({"r": r, "cur_price": cur_price, "hist": hist, "fp": fp, - "scen": scen, "outlook": outlook, "comp": comp}) - return original(r, cur_price, hist, fp, scen, outlook, comp) + def spy(*a, **kw): + calls.append({"args": a, "kwargs": dict(kw)}) + return original(*a, **kw) L._decide = spy return calls, original +# Positional slots in `_decide(r, cur_price, hist, fp, scen, outlook, comp, cover_days)`. +_SLOTS = ("r", "cur_price", "hist", "fp", "scen", "outlook", "comp", "cover_days") + + +def bound(call: dict) -> dict: + """A captured call as a name->value dict, whichever way the args were passed.""" + out = dict(zip(_SLOTS, call["args"])) + out.update(call["kwargs"]) + return out + + def with_comp_match(call: dict, rival: float | None): """The captured scenario grid, plus the `comp_match` candidate build_live_sku would add. @@ -101,10 +118,18 @@ def with_comp_match(call: dict, rival: float | None): "cover_days": round(inventory / max(units, 0.1)), } import pandas as pd + + # REPLACE, never append. build_live_sku already adds a `comp_match` row whenever competitor + # state was usable at pipeline time; concatenating a second one gave the frame a duplicate + # index, and `.loc["comp_match", "price"]` then returned a two-row Series instead of a + # price β€” which crashed the whole backtest on exactly the SKUs that had real competitor + # data, i.e. the ones it most needed to report on. + scen = scen[scen["scenario"] != "comp_match"] return pd.concat([scen, pd.DataFrame([extra])], ignore_index=True) -def rerun(call: dict, comp, *, rival: float | None = None) -> tuple[str, str, str, float]: +def rerun(call: dict, comp, *, rival: float | None = None + ) -> tuple[str, str, str, float, float]: """Re-run the captured input through the real cascade with a different comp state. Returns the SHIPPABLE price too, i.e. after the same guardrails build_live_sku applies: @@ -115,19 +140,36 @@ def rerun(call: dict, comp, *, rival: float | None = None) -> tuple[str, str, st scen = with_comp_match(call, rival) if rival else call["scen"] action, rec, _cons, _aggr, reasons, _root, objective = L._decide( call["r"], call["cur_price"], call["hist"], call["fp"], scen, - call["outlook"], comp, + call["outlook"], comp, call.get("cover_days"), ) cur = call["cur_price"] - target = float(scen.set_index("scenario").loc[rec, "price"]) + # `.loc[key]` on a duplicated index yields a Series, not a scalar. Take the first match + # explicitly so a malformed grid is a wrong number rather than a TypeError deep in a run. + hit = scen.loc[scen["scenario"] == rec, "price"] + if hit.empty: + raise KeyError(f"scenario {rec!r} missing from the grid: " + f"{sorted(scen['scenario'])}") + target = float(hit.iloc[0]) floor = L.price_floor(call["r"], cur).get("floor") or 0.0 if action == "Investigate": - shipped = cur + # Investigate HOLDS β€” it proposes no price, so it has no target to test against the + # floor. `None` says that explicitly. (The live engine behaves the same way and + # reports the breach in words: "selling $X BELOW the floor β€” but price is not the + # lever here".) Scoring the held price as a "target below break-even" turned every + # already-underpriced SKU into a fake invariant violation. + return action, ",".join(reasons), objective, round(cur, 2), None else: - shipped = max(target, floor) if floor else target + floored = max(target, floor) if floor else target + shipped = floored step = cur * L.MAX_STEP_PCT if abs(shipped - cur) > step + 0.005: shipped = cur + (step if shipped > cur else -step) - return action, ",".join(reasons), objective, round(shipped, 2) + # `floored` is the TARGET after the floor is applied but BEFORE the step cap. That is the + # number the break-even invariant is about. `shipped` can legitimately sit under the floor + # when a SKU is already priced below it: the cap only permits a 5% move, so the engine + # walks up over several cycles and says so ("moving $X now; on the way to $Y"). Judging + # the invariant on `shipped` reports that intended climb as a violation. + return action, ",".join(reasons), objective, round(shipped, 2), round(floored, 2) def state_from_disk(asin: str | None, our_price: float): @@ -159,15 +201,62 @@ def state_from_disk(asin: str | None, our_price: float): max_age_hours=rules.competitor_state_max_age_hours, now=NOW) -def cf(status: BuyBoxStatus, our_price: float, rival: float | None) -> CompetitiveState: +def cf(status: BuyBoxStatus, our_price: float, rival: float | None, + basis: str | None = None) -> CompetitiveState: """A COUNTERFACTUAL state β€” fresh by construction, so the age gate cannot mask the rule.""" + from pricing_agent.competitive_state import BASIS_SAME_ASIN + return CompetitiveState( status=status, our_price=our_price, buy_box_price=rival, competitor_min=rival, competitor_median=rival, rivals=1 if rival else 0, source="counterfactual", + basis=basis or BASIS_SAME_ASIN, as_of=NOW, reason="counterfactual injected by the verdict backtest", ) +@contextmanager +def old_policy(): + """Restore the cascade's PRE-CHANGE behaviour for the two policy fixes under test. + + Both are re-expressed through the switches the change itself introduced, rather than by + keeping a second copy of the old code around β€” so the "before" arm cannot drift from what + the "after" arm actually turned off. + + * sheet undercuts: `competitor_sheet_requires_corroboration` off -> any material + like-for-like undercut cuts, even on a healthy Buy-Box-winning SKU. + * ad spiral: the yield set narrowed back to BUYBOX_SUPPRESSED only -> AD_SPIRAL again + overrides NO_COST_DATA and LOW_STOCK. + """ + from config.settings import get_rules + + rules = get_rules() + prev_corr = rules.competitor_sheet_requires_corroboration + prev_yield = L.AD_SPIRAL_YIELDS_TO + rules.competitor_sheet_requires_corroboration = False + L.AD_SPIRAL_YIELDS_TO = frozenset({"BUYBOX_SUPPRESSED"}) + try: + yield + finally: + rules.competitor_sheet_requires_corroboration = prev_corr + L.AD_SPIRAL_YIELDS_TO = prev_yield + + +def spiral(call: dict, cover: int | None = None) -> dict: + """The captured call with an unrecoverable ad curve (and optionally a cover override). + + `ad_curve_unrecoverable` is a property of the SKU's own fitted ad model, so most real SKUs + do not have it and the AD_SPIRAL ordering could not otherwise be exercised on live data. + Everything else about the call stays real. + """ + r = call["r"].model_copy(update={"ad_curve_unrecoverable": True, + "contribution_per_dollar": 0.05, + "break_even_ad_curve": None}) + out = dict(call, r=r) + if cover is not None: + out["cover_days"] = cover + return out + + def main() -> int: ap = argparse.ArgumentParser() ap.add_argument("--skus", default=",".join(DEFAULT_SKUS)) @@ -195,7 +284,7 @@ def main() -> int: # Pair each captured call with its SKU by current price β€” build_live_sku calls _decide # exactly once per SKU. - by_price = {round(c["cur_price"], 4): c for c in calls} + by_price = {round(c["cur_price"], 4): c for c in map(bound, calls)} rows = [] for sku in skus: @@ -222,13 +311,40 @@ def main() -> int: "CF_UNDERCUT": rerun(call, cf(BuyBoxStatus.LOST_PRICE, cur, rival), rival=rival), "CF_SUPPRESSED": rerun(call, cf(BuyBoxStatus.SUPPRESSED, cur, None)), } - rows.append((sku, cur, comp, arms)) + + # ── the two POLICY changes under test, each old-vs-new on this SKU's real data ── + from pricing_agent.competitive_state import BASIS_SHEET + + sheet_won = cf(BuyBoxStatus.WON, cur, rival, basis=BASIS_SHEET) + low_stock = spiral(call, cover=20) + cases = { + # A rival BRAND is `--undercut` cheaper while we hold our own Buy Box. + "SHEET_UNDERCUT_WE_WIN": (call, sheet_won, rival), + # An unrecoverable ad curve on a SKU that is about to run out of stock. + "AD_SPIRAL_LOW_STOCK": (low_stock, None, None), + # ...and with the COGS fields empty, which is what makes the curve untrustworthy. + "AD_SPIRAL_NO_COST": (dict(spiral(call), fp={**call["fp"], "cost": 0.0, + "fba": 0.0}), None, None), + } + policy = {} + for name, (c, state_, riv) in cases.items(): + new = rerun(c, state_, rival=riv) + with old_policy(): + old = rerun(c, state_, rival=riv) + policy[name] = (old, new) + # Carry THIS SKU's floor and counterfactual rival on the row. They used to be read + # inside the reporting loop from `call` / `rival`, which by then were whatever the + # last iteration of THIS loop left behind β€” so every SKU printed the final SKU's + # break-even. On the run that found this, one SKU's floor showed as $14.12 while its + # real floor was $22.80. + floor = L.price_floor(call["r"], cur).get("floor") or 0.0 + rows.append((sku, cur, comp, arms, policy, floor, rival)) print("\n" + "=" * 100) print("VERDICT BACKTEST β€” competitor rules blind vs live, on real SKUs") print("=" * 100) - for sku, cur, comp, arms in rows: + for sku, cur, comp, arms, _policy, floor, rival in rows: print(f"\n{sku} current ${cur:.2f}") st = (f"{comp.status.value} via {comp.source}" if comp is not None else "none") usable = "usable" if (comp is not None and comp.usable) else ( @@ -238,22 +354,58 @@ def main() -> int: if comp is not None and comp.competitor_min: print(f" : cheapest rival ${comp.competitor_min:.2f} " f"({comp.rivals} rival offer(s))") - floor = L.price_floor(call["r"], cur).get("floor") or 0.0 print(f" break-even floor : ${floor:.2f}" - f" (counterfactual rival ${rival:.2f})") - for arm, (action, reasons, obj, shipped) in arms.items(): + + (" β€” ALREADY PRICED BELOW IT; the 5% cap means several cycles to climb out" + if floor and cur < floor - 0.005 else "") + + f" (counterfactual rival ${rival:.2f})") + for arm, (action, reasons, obj, shipped, target) in arms.items(): moved = "" if arms[arm] == arms["BLIND"] else " <<< CHANGED" tag = " (counterfactual)" if arm.startswith("CF_") else "" flag = "" - if floor and shipped < floor - 0.005: - flag = " *** BELOW FLOOR β€” INVARIANT VIOLATED ***" + if floor and target is not None and target < floor - 0.005: + flag = " *** TARGETS BELOW BREAK-EVEN β€” INVARIANT VIOLATED ***" + elif floor and shipped < floor - 0.005: + flag = (" (holding below the floor β€” price is not the lever)" + if target is None else + " (step-capped, still climbing to the floor)") print(f" {arm:14} {action:12} ${shipped:>7.2f} {reasons:42} " f"{obj}{moved}{tag}{flag}") + # ── POLICY DELTA: what the two cascade changes do, old vs new, on these SKUs ── print("\n" + "=" * 100) - real_changed = [s for s, _c, _st, a in rows if a["REAL"] != a["BLIND"]] - cf_u_changed = [s for s, _c, _st, a in rows if a["CF_UNDERCUT"] != a["BLIND"]] - cf_s_changed = [s for s, _c, _st, a in rows if a["CF_SUPPRESSED"] != a["BLIND"]] + print("POLICY DELTA β€” cascade BEFORE vs AFTER the two changes") + print("=" * 100) + CASE_NOTE = { + "SHEET_UNDERCUT_WE_WIN": + f"a rival BRAND is {args.undercut:.0%} cheaper, we HOLD our own Buy Box, and this " + f"SKU's real demand decides whether that is corroborated", + "AD_SPIRAL_LOW_STOCK": + "unrecoverable ad curve on a SKU with 20 days of cover (counterfactual curve)", + "AD_SPIRAL_NO_COST": + "unrecoverable ad curve with costPerUnit/fbaFee empty (counterfactual curve)", + } + deltas = {k: [] for k in CASE_NOTE} + for sku, cur, _comp, _arms, policy, _floor, _rival in rows: + print(f"\n{sku} current ${cur:.2f}") + for case, (old, new) in policy.items(): + changed = old != new + if changed: + deltas[case].append(sku) + print(f" {case}") + print(f" {CASE_NOTE[case]}") + print(f" before {old[0]:12} ${old[3]:>7.2f} {old[1]}") + print(f" after {new[0]:12} ${new[3]:>7.2f} {new[1]}" + + (" <<< CHANGED" if changed else " (unchanged)")) + + print("\n" + "-" * 100) + for case, skus_changed in deltas.items(): + print(f"{case:24} verdict changed on {len(skus_changed)}/{len(rows)} " + f"{skus_changed}") + + print("\n" + "=" * 100) + real_changed = [r[0] for r in rows if r[3]["REAL"] != r[3]["BLIND"]] + cf_u_changed = [r[0] for r in rows if r[3]["CF_UNDERCUT"] != r[3]["BLIND"]] + cf_s_changed = [r[0] for r in rows if r[3]["CF_SUPPRESSED"] != r[3]["BLIND"]] print(f"SKUs analysed : {len(rows)}") print(f"Verdicts changed by REAL competitor data : {len(real_changed)} " f"{real_changed}") @@ -263,15 +415,26 @@ def main() -> int: f"{cf_s_changed}") # The invariant that matters most: no arm, however cheap the counterfactual rival, may # ship a price under break-even. - violations = [] - for sku, cur, _comp, arms in rows: - floor = L.price_floor(by_price[round(cur, 4)]["r"], cur).get("floor") or 0.0 + # Measured on the TARGET, not the step-capped first move. A SKU already selling under its + # own ad-inclusive floor cannot be lifted over it in one 5% step, and reporting that + # deliberate multi-cycle climb as a breach buries any real one. Both are counted, and the + # climb is reported separately so it stays visible. + violations, climbing = [], [] + for sku, cur, _comp, arms, policy, floor, _rival in rows: if not floor: continue - for arm, (_action, _reasons, _obj, shipped) in arms.items(): - if shipped < floor - 0.005: - violations.append(f"{sku}/{arm} ${shipped:.2f} < ${floor:.2f}") - print(f"Prices shipped BELOW break-even (any arm) : {len(violations)} {violations}") + checks = dict(arms) + # The policy arms ship prices too, so they are held to the same invariant. + for case, (old, new) in policy.items(): + checks[f"{case}/before"], checks[f"{case}/after"] = old, new + for arm, (_action, _reasons, _obj, shipped, target) in checks.items(): + if target is not None and target < floor - 0.005: + violations.append(f"{sku}/{arm} targets ${target:.2f} < ${floor:.2f}") + elif shipped < floor - 0.005: + climbing.append(f"{sku}/{arm} ${shipped:.2f} -> ${floor:.2f}") + print(f"Targets BELOW break-even (real violations) : {len(violations)} {violations}") + print(f"Below floor but stepping UP toward it (by design): {len(set(climbing))} " + f"{sorted({c.split('/')[0] for c in climbing})}") print("\nA REAL delta of 0 is the fail-safe working, not the feature missing: with no " "usable\ncompetitor state every verdict is byte-identical to the competitor-blind " "one.\nThe counterfactual arms show the rules do fire once state IS usable.") diff --git a/src/pricing_agent/analyze.py b/src/pricing_agent/analyze.py index 64f4860..c5342e8 100644 --- a/src/pricing_agent/analyze.py +++ b/src/pricing_agent/analyze.py @@ -84,8 +84,14 @@ class AnalysisResult(BaseModel): # EVIDENCE (leads): what actually happened at each observed price price_performance: list = [] # monthly: price, units/day, ACTUAL profit/day price_bands: list = [] # pooled by price band + # The price customers ACTUALLY paid in the best band (`avg_price`), never the + # band's $0.50-rounded key β€” recommending the key puts a price on screen that + # was never charged, off by up to half a band width. best_observed_price: float | None = None best_observed_profit_day: float | None = None + # Days of real sample behind `best_observed_price`. Carried explicitly because + # re-deriving it downstream by matching a price back to its band is fragile. + best_observed_days: int | None = None actual_profit_per_unit: float | None = None unprofitable_months: int = 0 unmodeled_cost_gap: float | None = None # modelled net βˆ’ actual profit per unit @@ -477,15 +483,20 @@ def _elasticity_block(svc, sku: str, stack, rules, history=None) -> dict: take_home = stack.profit net_incl_ad = round(take_home - ad, 2) if best: - logger.info("evidence: best observed price $%.2f -> $%.0f/day actual profit " - "(%d/%d months unprofitable)", best["price_band"], - best["actual_profit_per_day"], n_bad, len(perf)) + logger.info("evidence: best observed price $%.2f (band $%.2f, %d days) -> $%.0f/day " + "actual profit (%d/%d months unprofitable)", best["avg_price"], + best["price_band"], best["days"], best["actual_profit_per_day"], + n_bad, len(perf)) out = { "price_performance": perf, "price_bands": bands, - "best_observed_price": (best["price_band"] if best else None), + # `avg_price`, not `price_band`: the band key is a $0.50-rounded bucket label, + # so quoting it recommends a price the SKU never actually ran (12.50 for a band + # that averaged 12.37). Every downstream consumer treats this as a TARGET PRICE. + "best_observed_price": (best["avg_price"] if best else None), "best_observed_profit_day": (best["actual_profit_per_day"] if best else None), + "best_observed_days": (best["days"] if best else None), "actual_profit_per_unit": actual_ppu, "unprofitable_months": n_bad, "unmodeled_cost_gap": reconcile(actual_ppu, net_incl_ad), diff --git a/src/pricing_agent/cosmos/service.py b/src/pricing_agent/cosmos/service.py index 4707650..75c92aa 100644 --- a/src/pricing_agent/cosmos/service.py +++ b/src/pricing_agent/cosmos/service.py @@ -59,21 +59,68 @@ class CosmosPricingService: # ── Product lookup ──────────────────────────────────────────────────── def get_product(self, sku: str) -> CosmosProduct | None: """Best-effort product enrichment (ASIN, cost, velocity). Never raises.""" + # `sku=` (a SKU filter), NOT `q=` (relevance search). `q` ranks by relevance across + # the whole catalogue and routinely does not return the SKU asked for at all: for + # UBMICROFIBERDUVETTWINWHITE the exact row sat on page 2 of 100-row pages, behind 100 + # unrelated products, so a 20-row `q` lookup never saw it. `sku=` puts it first. + # + # `sku=` is still a CONTAINS filter, so it is not sufficient on its own β€” see + # _exact_row. UBCFKMATTRESSPROTECTORTWIN88 returns three rows: the real one, the + # ...BOX variant, and a WAL... Walmart row whose "ASIN" (8946709597) is not even an + # ASIN. All three are marketplace AMAZON_USA, so only the exact SKU test separates + # them. try: body = self.client.get("/api/products", { - "q": sku, "page": 1, "size": 20, + "sku": sku, "page": 1, "size": 20, "marketplaces": _MARKETPLACES_FILTER, "targetCurrency": "USD", }) except CosmosApiError as e: logger.warning("product lookup failed for %s: %s", sku, e) return None page = Page.model_validate(body or {}) - for row in page.data: - if str(row.get("sku", "")).upper() == sku.upper(): - return CosmosProduct.model_validate(row) - # Fall back to the first row if COSMOS returned a fuzzy match set. - if page.data: - return CosmosProduct.model_validate(page.data[0]) + row = self._exact_row(page.data, sku, endpoint="/api/products") + return CosmosProduct.model_validate(row) if row is not None else None + + def _exact_row(self, rows: list[dict], sku: str, *, endpoint: str) -> dict | None: + """The row whose SKU matches `sku` EXACTLY, on this marketplace. Else None. + + `q` on /api/products and `skuPrefix` on /api/invp-insight are both RELEVANCE / + PREFIX searches, not equality filters β€” they happily return a set that does not + contain the SKU asked for. Both call sites used to fall back to `data[0]` "if + COSMOS returned a fuzzy match set", which silently bound one product's data to a + different product's SKU. + + Observed: `get_product("UBCFKFITTEDSHEETWHITECALKING")` β€” a SKU COSMOS does not + have at all ("Product not found" from the fee endpoint) β€” returned the row for + UBMICROFIBERGUSSETPILLOWWHITEQUEEN, ASIN B08DTH86Q2. Downstream, `analyze_price` + assigns `asin = product.asin` whenever INVP has none, so the competitive scrape + would have run against an unrelated listing and the dashboard would have shown + that listing's brand and marketplace under this SKU's name. + + The prefix case is the same bug with a likelier trigger: `skuPrefix=UBMICRO...TWIN` + matches every colour variant, and `data[0]` is then an arbitrary sibling whose + inventory and velocity drive the low-stock and excess-stock rules. + + A wrong number is worse than a blank one, so this fails safe: no exact match, no + record. Marketplace is checked too β€” CA/TEST/BOX variants share SKU codes with the + US catalogue, and a row from one must never answer for the other. + """ + want = sku.upper() + for row in rows or []: + if str(row.get("sku", "")).upper() != want: + continue + mkt = str(row.get("marketplace") or "").upper() + if mkt and mkt != self.marketplace.upper(): + logger.info("%s %s: skipping %s row (want %s)", + endpoint, sku, mkt, self.marketplace) + continue + return row + if rows: + logger.warning( + "%s %s: no EXACT match on %s in %d row(s) (got %s) β€” returning nothing " + "rather than another SKU's data", + endpoint, sku, self.marketplace, len(rows), + ", ".join(str(r.get("sku")) for r in rows[:3])) return None def list_skus(self, limit: int = 20, sku_prefix: str | None = None, @@ -274,8 +321,10 @@ class CosmosPricingService: logger.warning("invp lookup failed for %s: %s", sku, e) return None page = Page.model_validate(body or {}) - row = next((r for r in page.data if str(r.get("sku", "")).upper() == sku.upper()), - page.data[0] if page.data else None) + # `skuPrefix` is a PREFIX search: for UBMICROFIBERDUVETTWIN every colour variant + # comes back. Taking data[0] attached a sibling's inventory, cover days and 6-month + # velocity to this SKU β€” the inputs to LOW_STOCK and EXCESS_STOCK. See _exact_row. + row = self._exact_row(page.data, sku, endpoint="/api/invp-insight") if row is None: logger.info("invp %s β†’ no record", sku) return None diff --git a/tests/test_competitor_rules.py b/tests/test_competitor_rules.py index 8b00520..192cfaa 100644 --- a/tests/test_competitor_rules.py +++ b/tests/test_competitor_rules.py @@ -83,6 +83,17 @@ def scen(comp_match: float | None = None): return _scenarios(CUR, 10.0, 5000.0, FP, -1.3, None, extra) +def dropping() -> FakeResult: + """A SKU whose demand has materially fallen β€” 30d well under 70% of the 6-month mean. + + A LIKE-FOR-LIKE SHEET undercut only cuts a price when something we can independently + measure agrees the rival matters (see `competitor_sheet_requires_corroboration`). Tests + that are about WHICH RIVAL PRICE gets matched, rather than about that policy, use this so + they keep exercising the matching logic instead of silently testing the gate. + """ + return FakeResult(avg_30d=6.0, avg_6m=10.0) + + def state(status: BuyBoxStatus, *, rival: float | None = None, our: float = CUR, age_h: float = 0.0, rivals: int = 1, max_age_hours: float = 6.0 ) -> CompetitiveState: @@ -251,6 +262,58 @@ def test_suppression_is_not_relabelled_as_an_ad_problem(): assert verdict(state(BuyBoxStatus.SUPPRESSED), r=r)[2] == ["BUYBOX_SUPPRESSED"] +# ------------------------------------------------- what AD_SPIRAL may not override +# AD_SPIRAL is applied AFTER the cascade and overrides whatever fired, so anything it must +# not bury has to be named explicitly. These are the three that outrank it. +def test_ad_spiral_does_not_bury_missing_cost_data(): + """With costPerUnit/fbaFee at 0 the ad-curve maths is not even trustworthy. + + `fixed` is understated, which makes the "contribution <= 0" test EASIER to satisfy β€” so a + SKU whose real problem is an empty COGS field could be sent to the ad console instead of + to the data-entry field that actually fixes it. + """ + r = FakeResult(ad_curve_unrecoverable=True, contribution_per_dollar=0.05) + no_cost = {**FP, "cost": 0.0, "fba": 0.0} + action, _rec, _cons, _aggr, reasons, _root, _obj = _decide( + r, CUR, hist(), no_cost, scen(), None, None) + assert reasons == ["NO_COST_DATA"] + assert action == "Investigate" + + +def test_ad_spiral_does_not_bury_low_stock(): + """The only one of the three that changes a PRICE rather than a label. + + A shelf about to empty gets +5% to slow the burn, and that is right whatever the ad slope + is doing: those units sell regardless, so the only question is what we get for them. + Holding instead sells the last of the stock cheap. + """ + r = FakeResult(ad_curve_unrecoverable=True, contribution_per_dollar=0.05, cover_days=20) + action, rec, _cons, _aggr, reasons, _root, obj = _decide( + r, CUR, hist(), FP, scen(), None, None, 20) + assert "LOW_STOCK" in reasons + assert "AD_SPIRAL" not in reasons + assert (action, rec, obj) == ("Increase", "up_5", "low_stock_protection") + + +def test_ad_spiral_still_wins_over_everything_else(): + """The gate is narrow: three named findings, not a general demotion. + + Overstock is the mirror case of low stock and deliberately does NOT outrank it β€” cutting + price to clear stock is exactly the move that cannot work when ads eat the contribution. + """ + r = FakeResult(ad_curve_unrecoverable=True, contribution_per_dollar=0.05, cover_days=200) + _a, _rec, _c, _ag, reasons, _root, obj = _decide( + r, CUR, hist(), FP, scen(), None, None, 200) + assert reasons == ["AD_SPIRAL"] and obj == "fix_ad_efficiency" + + +def test_ad_spiral_yield_set_is_the_documented_one(): + """Guards the set itself β€” silently widening it would quietly demote the rule.""" + from dashboard.live_data import AD_SPIRAL_YIELDS_TO + + assert AD_SPIRAL_YIELDS_TO == {"NO_COST_DATA", "BUYBOX_SUPPRESSED", "LOW_STOCK"} + + # ============================================================ 3. UNDERCUT def test_material_undercut_recommends_a_decrease(): action, rec, reasons = verdict(state(BuyBoxStatus.LOST_PRICE, rival=22.0), @@ -664,8 +727,9 @@ def test_an_otherwise_identical_exact_match_does_trigger_it(tmp_path): st = s.state_for("UBTESTLINEKINGEXACT", our_price=None) assert st.competitor_min == 24.00 assert st.undercut_pct == pytest.approx(0.20) - action, rec, reasons = verdict(st, comp_match=24.00) - assert (action, rec, reasons) == ("Decrease", "comp_match", ["COMPETITOR_UNDERCUT"]) + action, rec, reasons = verdict(st, r=dropping(), comp_match=24.00) + assert (action, rec) == ("Decrease", "comp_match") + assert reasons == ["COMPETITOR_UNDERCUT", "VELOCITY_DROP"] def test_fuzzy_exclusion_is_per_rival_not_per_row(tmp_path): @@ -683,7 +747,9 @@ def test_fuzzy_exclusion_is_per_rival_not_per_row(tmp_path): assert st.competitor_min == 27.00 # NOT 20.00 assert st.undercut_pct == pytest.approx(0.10) # It still fires β€” off the exact rival, at the exact rival's price. - assert verdict(st, comp_match=27.00)[2] == ["COMPETITOR_UNDERCUT"] + action, rec, reasons = verdict(st, r=dropping(), comp_match=27.00) + assert (action, rec) == ("Decrease", "comp_match") + assert "COMPETITOR_UNDERCUT" in reasons def test_an_excluded_fuzzy_rival_is_still_reported_with_its_price(tmp_path): @@ -785,10 +851,11 @@ def test_stale_sheet_is_na_and_names_the_line_to_re_run(tmp_path): def test_sheet_basis_undercut_fires_without_a_buybox_loss(tmp_path): - """A rival BRAND's cheaper product is a real signal even while we hold our own Buy Box. + """A rival BRAND's cheaper product can cut a price WITHOUT us having lost the Buy Box. This is the difference the `basis` field exists to record: it is NOT a Buy Box loss, and - must never be reported as one. + must never be reported as one. What it now additionally needs is CORROBORATION β€” here, + demand that has actually fallen. See the uncorroborated case directly below. """ from pricing_agent.competitive_state import BASIS_SHEET @@ -796,6 +863,70 @@ def test_sheet_basis_undercut_fires_without_a_buybox_loss(tmp_path): st = s.state_for("UBTESTLINEQUEENWHITE", our_price=None) assert st.status is BuyBoxStatus.WON # we hold it, and it still fires assert st.basis == BASIS_SHEET + action, rec, reasons = verdict(st, r=dropping(), comp_match=24.99) + assert (action, rec) == ("Decrease", "comp_match") + assert reasons == ["COMPETITOR_UNDERCUT", "VELOCITY_DROP"] + + +def test_sheet_undercut_alone_does_not_cut_a_healthy_buybox_winning_sku(tmp_path): + """The margin leak this gate closes. + + A different brand being 13.8% cheaper is a market fact, not proof it is costing us + anything: we hold our own Buy Box and demand has not moved. Cutting there spends certain + margin against an unmeasured threat β€” and because the competitor branch sits ABOVE + PROFIT_OPTIMAL, it could overrule an elasticity fit that wanted a RAISE. + """ + s = _sheet(tmp_path) + st = s.state_for("UBTESTLINEQUEENWHITE", our_price=None) + assert st.status is BuyBoxStatus.WON + assert st.undercut_pct == pytest.approx(0.1379, abs=1e-3) # materially cheaper + # FakeResult's demand is flat (30d == 6m), so nothing corroborates the gap. + assert verdict(st, comp_match=24.99) == verdict(None) + + +def test_sheet_undercut_still_cuts_when_we_are_not_winning_the_buybox(tmp_path): + """The second form of corroboration: we do not hold the Buy Box. + + Losing it is itself the independent evidence that the cheaper rival is costing us + something, so no velocity drop is required as well. + """ + from pricing_agent.competitive_state import BASIS_SHEET + + s = _sheet(tmp_path) + st = s.state_for("UBTESTLINEQUEENWHITE", our_price=None) + st.status = BuyBoxStatus.LOST_PRICE + assert st.basis == BASIS_SHEET + action, rec, reasons = verdict(st, comp_match=24.99) + assert (action, rec, reasons) == ("Decrease", "comp_match", ["COMPETITOR_UNDERCUT"]) + + +def test_uncorroborated_sheet_undercut_is_reported_even_though_it_does_not_cut(tmp_path): + """A rival we deliberately did not chase must not look like missing data. + + "No competitor data" and "a rival is 13.8% cheaper and we chose not to react" are very + different states, and only one of them is worth re-running a scrape for. + """ + s = _sheet(tmp_path) + st = s.state_for("UBTESTLINEQUEENWHITE", our_price=None) + _a, _rec, _cons, _aggr, _reasons, root, _obj = _decide( + FakeResult(), CUR, hist(), FP, scen(24.99), None, st) + note = next((v for k, v in root if k == "Competitor undercut not acted on"), None) + assert note is not None + assert "13.8% cheaper" in note and "$24.99" in note + assert "no price cut" in note + + +def test_corroboration_gate_can_be_switched_off_in_config(tmp_path, monkeypatch): + """It is a policy judgement, so it is tunable without a deploy β€” like its neighbours.""" + from config.settings import get_rules + + s = _sheet(tmp_path) + st = s.state_for("UBTESTLINEQUEENWHITE", our_price=None) + assert verdict(st, comp_match=24.99) == verdict(None) # gated on + + rules = get_rules() + monkeypatch.setattr(rules, "competitor_sheet_requires_corroboration", False) + monkeypatch.setattr("config.settings.get_rules", lambda: rules) action, rec, reasons = verdict(st, comp_match=24.99) assert (action, rec, reasons) == ("Decrease", "comp_match", ["COMPETITOR_UNDERCUT"]) diff --git a/tests/test_cosmos.py b/tests/test_cosmos.py index b033caa..3733564 100644 --- a/tests/test_cosmos.py +++ b/tests/test_cosmos.py @@ -54,3 +54,99 @@ def test_stack_from_quote_matches_cosmos_net(): cost_base = 8.59 + 8.97 + 0.5 + 0.11 assert round(stack.break_even_price, 2) == round(cost_base / (1 - 0.15), 2) assert round(stack.map_floor, 2) == round(cost_base / (1 - 0.25), 2) + + +# ---------------------------------------------------------------- exact SKU joins +# COSMOS's /api/products and /api/invp-insight filters are CONTAINS/relevance searches, +# not equality. Both lookups used to fall back to `data[0]` "if COSMOS returned a fuzzy +# match set", which bound one product's ASIN, cost, brand and inventory to a different +# product's SKU. Observed live: get_product("UBCFKFITTEDSHEETWHITECALKING") β€” a SKU COSMOS +# does not have β€” returned UBMICROFIBERGUSSETPILLOWWHITEQUEEN / B08DTH86Q2. +class RowClient: + """Returns a fixed row set, and records the params it was asked for.""" + + def __init__(self, rows): + self.rows, self.calls = rows, [] + + def get(self, path, params=None): + self.calls.append((path, dict(params or {}))) + return {"data": self.rows, "page": 1, "size": 20, "total": len(self.rows)} + + +# The real response for sku=UBCFKMATTRESSPROTECTORTWIN88: the SKU itself, a Walmart row +# whose "ASIN" is not an ASIN, and the BOX variant. All three are AMAZON_USA, so only the +# exact SKU test tells them apart. +PROTECTOR_ROWS = [ + {"sku": "UBCFKMATTRESSPROTECTORTWIN88", "marketplace": "AMAZON_USA", + "asin": "B00MRH9NCK", "status": "Registered", "cost": 4.11}, + {"sku": "WALUBCFKMATTRESSPROTECTORTWIN88", "marketplace": "AMAZON_USA", + "asin": "8946709597", "cost": 9.99}, + {"sku": "UBCFKMATTRESSPROTECTORTWIN88BOX", "marketplace": "AMAZON_USA", + "asin": "B09K7HXJ4M", "cost": 7.77}, +] + + +def test_get_product_picks_the_exact_sku_not_a_box_or_walmart_variant(): + c = RowClient(PROTECTOR_ROWS) + p = CosmosPricingService(c).get_product("UBCFKMATTRESSPROTECTORTWIN88") + assert p is not None + assert p.sku == "UBCFKMATTRESSPROTECTORTWIN88" + assert p.asin == "B00MRH9NCK" # not 8946709597, not B09K7HXJ4M + assert p.cost == 4.11 # not another variant's COGS + + +def test_get_product_queries_the_sku_filter_not_the_relevance_search(): + """`q` ranks across the whole catalogue and often omits the SKU entirely.""" + c = RowClient(PROTECTOR_ROWS) + CosmosPricingService(c).get_product("UBCFKMATTRESSPROTECTORTWIN88") + _path, params = c.calls[0] + assert params.get("sku") == "UBCFKMATTRESSPROTECTORTWIN88" + assert "q" not in params + + +def test_get_product_returns_none_rather_than_another_skus_row(): + """The live failure: a SKU COSMOS does not carry at all.""" + c = RowClient([ + {"sku": "UBMICROFIBERGUSSETPILLOWWHITEQUEEN", "marketplace": "AMAZON_USA", + "asin": "B08DTH86Q2", "cost": 8.59}, + {"sku": "UBPILLOWSQUARECOTTONCOVER18X18", "marketplace": "AMAZON_USA", + "asin": "B07XYZ1234"}, + ]) + assert CosmosPricingService(c).get_product("UBCFKFITTEDSHEETWHITECALKING") is None + + +def test_get_product_rejects_a_foreign_marketplace_row(): + c = RowClient([{"sku": "UBTEST", "marketplace": "AMAZON_CA", "asin": "BCA0000001"}]) + svc = CosmosPricingService(c, marketplace="AMAZON_USA") + assert svc.get_product("UBTEST") is None + # ...and accepts it when that IS the marketplace we asked about. + assert CosmosPricingService(c, marketplace="AMAZON_CA").get_product("UBTEST") is not None + + +def test_get_invp_does_not_take_a_sibling_variants_inventory(): + """`skuPrefix` matches every colour/size variant; data[0] was an arbitrary sibling. + + Inventory, cover days and 6-month velocity feed LOW_STOCK and EXCESS_STOCK, so a + sibling's numbers here move real prices on the wrong SKU. + """ + def _row(sku, inv, cover): + # cover_days is derived from the EARLIEST dateMap snapshot, not a top-level field. + return {"sku": sku, "marketplace": "AMAZON_USA", "inventory": inv, + "averageSale6Months": inv / 100.0, + "dateMap": {"01/05/2026": {"dataDate": "01/05/2026", "inventory": inv, + "coverDays": cover}}} + + rows = [_row("UBMICROFIBERDUVETTWINGREY", 12, 3), + _row("UBMICROFIBERDUVETTWINWHITE", 4000, 78)] + invp = CosmosPricingService(RowClient(rows)).get_invp("UBMICROFIBERDUVETTWINWHITE") + assert invp is not None + assert invp.sku == "UBMICROFIBERDUVETTWINWHITE" + # Not the GREY sibling's 12 units / 3 days β€” which would read as a stockout emergency + # and fire LOW_STOCK on a SKU sitting on 4,000 units. + assert (invp.inventory, invp.cover_days, invp.avg_6m) == (4000, 78, 40.0) + + +def test_get_invp_returns_none_when_only_siblings_come_back(): + rows = [{"sku": "UBMICROFIBERDUVETTWINGREY", "marketplace": "AMAZON_USA", + "inventory": 12}] + assert CosmosPricingService(RowClient(rows)).get_invp("UBMICROFIBERDUVETTWINWHITE") is None diff --git a/tests/test_pricing_safety.py b/tests/test_pricing_safety.py index f426126..c494ede 100644 --- a/tests/test_pricing_safety.py +++ b/tests/test_pricing_safety.py @@ -288,3 +288,205 @@ def test_ladder_is_always_ordered(action, rec, cons, aggr): def test_ladder_untouched_for_hold_actions(): assert _order_ladder("Maintain", "current", "current", "up_2", LADDER_SCEN, 26.07) == ("current", "up_2") + + +# ------------------------------------------------- best-observed price provenance +# The band key is a $0.50-rounded BUCKET LABEL; `avg_price` is what customers really +# paid. Publishing the key as `best_observed_price` recommends a price the SKU never +# ran, and re-deriving the sample size by matching that key against each band's +# `avg_price` never matched β€” so the recommendation rationale claimed "0 days" of +# evidence for a price band with weeks behind it. +class _EvidenceResult: + """The subset of AnalysisResult that `_evidence_target` reads.""" + + def __init__(self, bop, bday, bands, days=None): + self.best_observed_price = bop + self.best_observed_profit_day = bday + self.best_observed_days = days + self.price_bands = bands + + +def _hist_days(price, units, profit, ad, n, month): + class _D: + def __init__(s, date, sale_price, units, profit, marketing_cost): + s.date, s.sale_price, s.units = date, sale_price, units + s.profit, s.marketing_cost = profit, marketing_cost + return [_D(f"{month:02d}/{i + 1:02d}/2025", price, units, profit, ad) + for i in range(n)] + + +class _Stack: + """The fee stack `_elasticity_block` reads β€” a cheap SKU, comfortably profitable.""" + + selling_price = 12.37 + referral_pct = 0.15 + returns_reserve = 0.25 + landed_cost = 4.00 + fba_fee = 3.00 + storage_alloc = 0.25 + profit = 2.00 + + +def test_best_observed_price_is_what_customers_paid_not_the_band_key(): + from config.settings import get_rules + from pricing_agent.analyze import _elasticity_block + from pricing_agent.performance import best_observed, price_band_performance + + # Six months, two price levels. $12.37 rounds into the $12.50 band and wins on + # actual profit/day; $13.90 rounds into the $14.00 band. + hist = [] + for month in (1, 2, 3): + hist += _hist_days(12.37, 10, 30.0, 5.0, 20, month) + for month in (4, 5, 6): + hist += _hist_days(13.90, 8, 20.0, 5.0, 20, month) + + bands = price_band_performance(hist, band=0.50, min_days=7) + best = best_observed(bands) + assert best["price_band"] == 12.5 # the bucket label + assert best["avg_price"] == 12.37 # the price actually charged + + out = _elasticity_block(None, "TESTSKU", _Stack(), get_rules(), history=hist) + # The published TARGET must be the charged price, never the rounded bucket label. + assert out["best_observed_price"] == 12.37 + assert out["best_observed_price"] != best["price_band"] + # ...and it must carry its own sample size, so nothing has to re-derive it. + assert out["best_observed_days"] == 60 + + +def _evidence_from_history(): + """Run the real analysis path, then hand its output to the real decision helper. + + Feeding `_evidence_target` a hand-picked price would test nothing: the bug was that + `_elasticity_block` published the BAND KEY, which `_evidence_target` then could not + match back to any band. Only the two together reproduce it. + """ + from config.settings import get_rules + from pricing_agent.analyze import _elasticity_block + + hist = [] + for month in (1, 2, 3): + hist += _hist_days(12.37, 10, 30.0, 5.0, 20, month) + for month in (4, 5, 6): + hist += _hist_days(15.90, 6, 6.0, 0.9, 20, month) + return _elasticity_block(None, "TESTSKU", _Stack(), get_rules(), history=hist) + + +def test_evidence_rationale_quotes_the_real_sample_size(): + from dashboard.live_data import _evidence_target + + out = _evidence_from_history() + r = _EvidenceResult(bop=out["best_observed_price"], + bday=out["best_observed_profit_day"], + bands=out["price_bands"], + days=out["best_observed_days"]) + got = _evidence_target(r, cur_price=15.90, keys={"best_observed"}) + assert got is not None + price, why, _gain = got + assert price == 12.37 # the charged price, not the $12.50 band + assert "over 60 days" in why + assert "over 0 days" not in why + + +def test_evidence_sample_size_falls_back_to_the_bands_when_field_absent(): + from dashboard.live_data import _evidence_target + + out = _evidence_from_history() + # Older callers that never set the field must still resolve the sample size β€” + # which works now only because `best_observed_price` IS a band's `avg_price`. + r = _EvidenceResult(bop=out["best_observed_price"], + bday=out["best_observed_profit_day"], + bands=out["price_bands"], days=None) + _price, why, _gain = _evidence_target(r, cur_price=15.90, keys={"best_observed"}) + assert "over 60 days" in why + + +# ------------------------------------------------- elasticity used for projections +# `elasticity_actionable` gated the DECISION but not the scenario projections, which +# took the raw point estimate whenever one existed. A CI spanning zero β€” or worse, a +# POSITIVE slope β€” then drove the 30-day impact tile, the portfolio opportunity total +# and the queue sort. A positive slope projects that raising price sells MORE units. +def _fit(elasticity, actionable): + return {"elasticity": elasticity, "actionable": actionable, "r2": 0.1, + "confidence": "none" if not actionable else "high", "why": "test"} + + +from dashboard.live_data import projection_elasticity as _project_with + + +@pytest.mark.parametrize("fit,expect_fitted", [ + (_fit(-1.8, True), True), # usable fit β€” its own slope + (_fit(-0.15, False), False), # CI spans zero + (_fit(0.40, False), False), # POSITIVE β€” would say a raise sells more + (None, False), # no fit at all + ({}, False), +]) +def test_projections_only_use_a_statistically_usable_elasticity(fit, expect_fitted): + from dashboard.live_data import FALLBACK_ELASTICITY + + el, fitted = _project_with(fit) + assert fitted is expect_fitted + assert el < 0, "a projection must never assume raising price sells more" + if not expect_fitted: + assert el == FALLBACK_ELASTICITY + + +def test_positive_elasticity_never_projects_volume_growth_on_a_raise(): + from dashboard.live_data import _scenarios + + fp = {"cost": 8.0, "fba": 4.0, "referral_pct": 0.15, "returns_pct": 0.02, + "variable": 0.5} + el, _ = _project_with(_fit(0.40, False)) + scen = _scenarios(25.0, 10.0, 5000.0, fp, el, None, None) + by = dict(zip(scen["scenario"], scen["units_day"])) + assert by["up_5"] < by["current"] < by["down_5"] + + +# ------------------------------------------------- economics quoted beside a headline +# When a move is step-capped, `rec_key` names the DESTINATION rung while `rec_price` is +# what we actually recommend today. Quoting the destination's units/profit/margin beside +# today's price promises what the move does not deliver. Observed live on +# UBMICROFIBERDUVETTWINWHITE: recommend $17.94 (67 units/day, 6.5% margin) on the way to +# $19.90 (58 units/day, 12.8%) β€” the card showed the $19.90 figures under a $17.94 header. +def _econ_row(key, price, units, net, margin): + return {"key": key, "price": price, "units_day": units, "net_30d": net, + "net_margin_pct": margin} + + +STEP_CAPPED = { + "rec_key": "best_observed", # the $19.90 destination the cascade chose + "rec_price": 17.94, # what the 5% step cap actually permits today + "rec_units_day": 0, "units_day": 0, "rec_profit_30d": 0, "profit_30d": 0, + "scen_econ": [ + _econ_row("current", 18.30, 65, 2482, 6.9), + _econ_row("recommended", 17.94, 67, 2322, 6.5), + _econ_row("best_observed", 19.90, 58, 4462, 12.8), + ], +} + + +def test_card_quotes_the_price_it_recommends_not_the_capped_destination(): + from app import econ_rows + + cur, rec = econ_rows(STEP_CAPPED) + assert cur["price"] == 18.30 + assert rec["price"] == 17.94, "must follow rec_price, not the rec_key rung" + assert (rec["units_day"], rec["net_30d"], rec["net_margin_pct"]) == (67, 2322, 6.5) + # The destination's much rosier figures must NOT be what gets shown. + assert rec["net_margin_pct"] != 12.8 + + +def test_uncapped_move_still_resolves_to_its_rung(): + from app import econ_rows + + d = dict(STEP_CAPPED, rec_price=19.90) + _cur, rec = econ_rows(d) + assert rec["price"] == 19.90 and rec["net_margin_pct"] == 12.8 + + +def test_econ_rows_degrades_to_empty_dicts_without_scenario_economics(): + from app import econ_rows + + cur, rec = econ_rows({"rec_key": "x", "rec_price": 10.0, "scen_econ": []}) + assert cur == {} and rec == {} + # Callers use .get(..., fallback), so empty must not raise. + assert cur.get("units_day", 42) == 42
Product DescriptionDaily Sale
{cfg["title"]}
' - f'
{cfg["sku"]} Β· {asin_html}{reviews}
' + f'
{esc_html(cfg["title"])}
' + f'
{esc_html(cfg["sku"])} Β· {asin_html}{reviews}
' f'
πŸ“¦ {on_hand:,.0f} on hand Β· 🚚 {total_inbound:,.0f} inbound' f'{f" (+{later_inbound:,.0f} later)" if later_inbound else ""} Β· ' f'class {cfg["inv_class"].capitalize()}