Talent tab: LinkedIn sourcing per job via Apify
New backend/talent/ module (runs + profiles tables, RBAC talent.* seed, frontend-polled run lifecycle) and frontend Talent screen: job + location pickers, paid-search confirm, big loader, ranked profile cards with match ring, profile detail modal, show-10-then-more, next-page re-search. Sourcing quality: current-title facet + skills-only keyword query, experience-range facet, thin-result broadening ladder, Utopia Brands/Deals current employees excluded actor-side and server-side, graded 0-100 relevance score (title phrase over keyword stuffing). Also narrows the overbroad **_**_**.py gitignore rule to alembic versions; it was silently swallowing app/tests __init__.py files, the engine smoke script, and the new talent plugin tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>pull/8/head
parent
202d05c4f9
commit
e1bd8526f3
|
|
@ -55,4 +55,6 @@ node_modules/
|
|||
frontend/dist/
|
||||
|
||||
**.pdf
|
||||
**_**_**.py
|
||||
# Per-machine alembic autogen revisions only — the old bare `**_**_**.py`
|
||||
# also swallowed any module with two underscores (e.g. test_talent_plugins.py).
|
||||
backend/migrations/versions/**_**_**.py
|
||||
|
|
@ -0,0 +1,5 @@
|
|||
"""Bulk ATS scoring engine."""
|
||||
|
||||
__all__ = ["__version__"]
|
||||
|
||||
__version__ = "0.1.0"
|
||||
|
|
@ -31,6 +31,26 @@ BUFFER_API=
|
|||
BUFFER_API_URL=https://api.buffer.com
|
||||
BUFFER_CHANNEL_ID=
|
||||
|
||||
# Talent sourcing via Apify (talent/). Token from console.apify.com → Settings →
|
||||
# API & Integrations. APIFY_TOKEN is honoured as a fallback name for the token.
|
||||
APIFY_API_TOKEN=
|
||||
APIFY_API_BASE=https://api.apify.com/v2
|
||||
APIFY_ACTOR_ID=harvestapi~linkedin-profile-search
|
||||
# Hard per-run cap; client requests are clamped to it. "Full" mode costs
|
||||
# $0.10 per search page + $0.004 per profile (~$0.20 for a 25-profile run).
|
||||
APIFY_MAX_RESULTS=25
|
||||
# Server-side spend ceiling per run (Apify maxTotalChargeUsd; minimum $0.10).
|
||||
APIFY_MAX_COST_USD=1.0
|
||||
# Own companies whose CURRENT employees must never appear in sourced results.
|
||||
# Names feed the always-on server-side filter (case-insensitive substring);
|
||||
# URLs feed the actor's excludeCurrentCompanies filter (full LinkedIn company
|
||||
# URLs) so those profiles are not even scraped. Comma-separated.
|
||||
APIFY_EXCLUDE_COMPANIES=Utopia Brands,Utopia Deals
|
||||
APIFY_EXCLUDE_COMPANY_URLS=https://www.linkedin.com/company/utopiadeals,https://www.linkedin.com/company/utopia-brands-usa,https://www.linkedin.com/company/utopiabrands
|
||||
# Short | Full | Full + email search
|
||||
APIFY_PROFILE_MODE=Full
|
||||
APIFY_TIMEOUT=30
|
||||
|
||||
OPENAI_API_KEY=
|
||||
OPENAI_MODEL=gpt-5.4-mini
|
||||
# Blank omits the parameter, for reasoning models that reject it.
|
||||
|
|
|
|||
|
|
@ -261,8 +261,8 @@ Additional rules that matter when you edit this code:
|
|||
|
||||
### `users/`
|
||||
Signup, login, refresh, CRUD, role assignment, and the RBAC machinery every other domain
|
||||
depends on. `users/permissions.py` defines the full `PermissionTag` vocabulary (13 modules ×
|
||||
8 actions = 104 tags) and the `require_permission(...)` dependency. A startup assertion
|
||||
depends on. `users/permissions.py` defines the full `PermissionTag` vocabulary (15 modules ×
|
||||
8 actions = 120 tags) and the `require_permission(...)` dependency. A startup assertion
|
||||
(`_assert_vocabulary_complete`) fails loudly if the tag list ever drifts from
|
||||
`PermissionModule × PermissionAction`.
|
||||
|
||||
|
|
@ -332,6 +332,19 @@ reads:
|
|||
mints a `type=reset` JWT carrying the code row id (`crid`), which is the only thing that
|
||||
authorises the new-password call.
|
||||
|
||||
### `talent/`
|
||||
LinkedIn talent sourcing via Apify. `POST /talent/runs/start` launches one paid actor run
|
||||
(default actor: HarvestAPI's no-cookie `linkedin-profile-search`) with a search query built
|
||||
deterministically from the job's title, requirements and location. There is no worker: the
|
||||
frontend polls `GET /talent/runs/status`, and the first poll that sees the run `SUCCEEDED`
|
||||
fetches the dataset and upserts `talent_profiles` in that same request — idempotent, so a
|
||||
closed tab loses nothing. Profiles are deduped per job by normalized LinkedIn URL
|
||||
(`uq_talent_profiles_job_url`); re-runs refresh fields but never resurrect a dismissed
|
||||
(`is_deleted`) profile. The raw dataset item is kept verbatim in `talent_profiles.raw`
|
||||
because item shapes vary per actor. A run is refused with 409 while another is active for
|
||||
the same job, and `APIFY_MAX_COST_USD` is passed as `maxTotalChargeUsd` so Apify enforces
|
||||
the spend ceiling server-side.
|
||||
|
||||
### `agent/`
|
||||
LangGraph state machine — see [The matching agent](#the-matching-agent).
|
||||
|
||||
|
|
@ -827,6 +840,7 @@ The engine reads its own settings through `app.core.config.get_settings()`, from
|
|||
| **Email API** (a Microsoft Graph proxy) | `inbox/` | `GET {EMAIL_URL}/emails`, `GET {EMAIL_URL}/emails/{id}`, `GET {EMAIL_URL}/sync/read-status`, `GET {EMAIL_URL}/sync/read-status/message/{id}` — Bearer `EMAIL_API_TOKEN` |
|
||||
| **Teams Mail API** | `notifications/`, `forget_password/` | multipart POST to `TEAMS_MAIL_API_URL`; success is HTTP **202**, anything else raises |
|
||||
| **Buffer** | `job/job_post/` | GraphQL against `BUFFER_API_URL` — `createPost` mutation, `account { organizations }` and `channels` queries |
|
||||
| **Apify** | `talent/` | REST against `APIFY_API_BASE` — `POST /acts/{id}/runs` (with `maxTotalChargeUsd`), `GET /actor-runs/{id}`, `GET /datasets/{id}/items` — Bearer `APIFY_API_TOKEN` |
|
||||
| **OpenAI** | `agent/`, `llm_setup.py` | Chat Completions with `response_format: json_object` |
|
||||
|
||||
Attachments are written to `backend/inbox/decoded_attachments/`. In Docker this directory is
|
||||
|
|
@ -898,6 +912,20 @@ own keys with `os.getenv`.
|
|||
| `BUFFER_API_URL` | `https://api.buffer.com` |
|
||||
| `BUFFER_CHANNEL_ID` | — (fallback channel) |
|
||||
|
||||
### Apify (talent sourcing)
|
||||
|
||||
| Variable | Default | Notes |
|
||||
|---|---|---|
|
||||
| `APIFY_API_TOKEN` | — | API token from console.apify.com → Settings → API & Integrations; `APIFY_TOKEN` accepted as a fallback name |
|
||||
| `APIFY_API_BASE` | `https://api.apify.com/v2` | |
|
||||
| `APIFY_ACTOR_ID` | `harvestapi~linkedin-profile-search` | `user~actor` form, as used in URL paths |
|
||||
| `APIFY_MAX_RESULTS` | `25` | Hard per-run profile cap; client requests are clamped to it |
|
||||
| `APIFY_PROFILE_MODE` | `Full` | `Short` \| `Full` \| `Full + email search` — `Full` is $0.10/search page + $0.004/profile (~$0.20 per 25-profile run) |
|
||||
| `APIFY_MAX_COST_USD` | `1.0` | Sent as `maxTotalChargeUsd`; Apify's minimum is $0.10 |
|
||||
| `APIFY_TIMEOUT` | `30` | Per-request httpx timeout, seconds |
|
||||
| `APIFY_EXCLUDE_COMPANIES` | `Utopia Brands,Utopia Deals` | Own companies: current employees are filtered out server-side before profiles are stored (case-insensitive substring on current company, headline fallback) |
|
||||
| `APIFY_EXCLUDE_COMPANY_URLS` | the Utopia Deals / Utopia Brands USA / Utopia Brands Pakistan pages | Full LinkedIn company URLs for the actor's `excludeCurrentCompanies` filter — stops those profiles being scraped (and billed) at all |
|
||||
|
||||
### OpenAI
|
||||
|
||||
| Variable | Default |
|
||||
|
|
|
|||
|
|
@ -412,11 +412,15 @@ async def fetch_job_posts(
|
|||
skip: int = Query(0, ge=0),
|
||||
ids: str | None = Query(None),
|
||||
active_only: bool = Query(True),
|
||||
# Either job-board or candidate viewers may list jobs — recruiters scoring
|
||||
# CVs need a job to score against (CV Import picker).
|
||||
# Job-board, candidate, or talent viewers may list jobs — recruiters scoring
|
||||
# CVs need a job to score against (CV Import picker), and talent sourcing
|
||||
# needs the same picker to choose which job to source for.
|
||||
current_user: dict = Depends(
|
||||
require_permission(
|
||||
PermissionTag.JOB_BOARD_VIEW, PermissionTag.CANDIDATES_VIEW, require_all=False
|
||||
PermissionTag.JOB_BOARD_VIEW,
|
||||
PermissionTag.CANDIDATES_VIEW,
|
||||
PermissionTag.TALENT_VIEW,
|
||||
require_all=False,
|
||||
)
|
||||
),
|
||||
session: AsyncSession = Depends(get_session),
|
||||
|
|
|
|||
|
|
@ -17,6 +17,7 @@ from assessments.app import router as assessments_router
|
|||
from org_settings.app import router as org_settings_router
|
||||
from saved_search.app import router as saved_search_router
|
||||
from search.app import router as search_router
|
||||
from talent.app import router as talent_router
|
||||
|
||||
logging.basicConfig(level=logging.INFO,format="%(levelname)-8s %(name)s: %(message)s")
|
||||
logger=logging.getLogger("main")
|
||||
|
|
@ -100,3 +101,4 @@ app.include_router(assessments_router)
|
|||
app.include_router(org_settings_router)
|
||||
app.include_router(saved_search_router)
|
||||
app.include_router(search_router)
|
||||
app.include_router(talent_router)
|
||||
|
|
|
|||
|
|
@ -0,0 +1,66 @@
|
|||
-- 007_talent_rbac.sql
|
||||
-- Manual one-shot: the `talent` permission module (8 tags), a `talent_sourcing`
|
||||
-- bundle holding them, and the bundle attached to the staff roles that source
|
||||
-- candidates. Mirrors 004's idempotent pattern; applied automatically at startup
|
||||
-- by alembic_setup.run_manual_sql() and recorded in manual_migrations.
|
||||
--
|
||||
-- The all_access bundle is a fixed id list seeded before this module existed,
|
||||
-- so system_administrator gets talent access through THIS bundle, not that one.
|
||||
-- Users must log in again after this applies — permissions are resolved from
|
||||
-- the DB per request, but the frontend caches the list from /users/me.
|
||||
|
||||
-- =============================================================================
|
||||
-- 1. The 8 talent.* permission tags
|
||||
-- =============================================================================
|
||||
INSERT INTO app.permission_tags
|
||||
(tag_name, module, action, description, created_at, updated_at, is_active, is_deleted)
|
||||
VALUES
|
||||
('talent.view', 'talent', 'view', NULL, NOW(), NOW(), true, false),
|
||||
('talent.create', 'talent', 'create', NULL, NOW(), NOW(), true, false),
|
||||
('talent.edit', 'talent', 'edit', NULL, NOW(), NOW(), true, false),
|
||||
('talent.delete', 'talent', 'delete', NULL, NOW(), NOW(), true, false),
|
||||
('talent.approve', 'talent', 'approve', NULL, NOW(), NOW(), true, false),
|
||||
('talent.export', 'talent', 'export', NULL, NOW(), NOW(), true, false),
|
||||
('talent.manage', 'talent', 'manage', NULL, NOW(), NOW(), true, false),
|
||||
('talent.configure', 'talent', 'configure', NULL, NOW(), NOW(), true, false)
|
||||
ON CONFLICT (tag_name) DO NOTHING;
|
||||
|
||||
-- =============================================================================
|
||||
-- 2. Bundle holding all eight talent tags
|
||||
-- =============================================================================
|
||||
INSERT INTO app.permissions (name, description, permission_tags, is_system, created_at, updated_at, is_active, is_deleted)
|
||||
SELECT
|
||||
'talent_sourcing',
|
||||
'LinkedIn talent sourcing: run Apify searches and view sourced profiles',
|
||||
(
|
||||
SELECT COALESCE(jsonb_agg(id ORDER BY id), '[]'::jsonb)
|
||||
FROM app.permission_tags
|
||||
WHERE is_deleted = false
|
||||
AND module = 'talent'
|
||||
),
|
||||
true,
|
||||
NOW(),
|
||||
NOW(),
|
||||
true,
|
||||
false
|
||||
WHERE NOT EXISTS (
|
||||
SELECT 1 FROM app.permissions WHERE name = 'talent_sourcing'
|
||||
);
|
||||
|
||||
-- =============================================================================
|
||||
-- 3. Attach the bundle to the staff roles (idempotent; same role list as 004)
|
||||
-- =============================================================================
|
||||
UPDATE app.roles r
|
||||
SET permissions = COALESCE(r.permissions, '[]'::jsonb) || jsonb_build_array(p.id),
|
||||
updated_at = NOW()
|
||||
FROM app.permissions p
|
||||
WHERE p.name = 'talent_sourcing'
|
||||
AND r.role_name IN (
|
||||
'system_administrator',
|
||||
'hr_administrator',
|
||||
'recruiter',
|
||||
'hiring_manager',
|
||||
'department_head',
|
||||
'ceo'
|
||||
)
|
||||
AND NOT (COALESCE(r.permissions, '[]'::jsonb) @> jsonb_build_array(p.id));
|
||||
|
|
@ -0,0 +1,118 @@
|
|||
from fastapi import APIRouter, Depends, HTTPException, Query
|
||||
from fastapi.responses import JSONResponse
|
||||
from pydantic import BaseModel
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from db_setup import get_session
|
||||
from talent.views import Talent
|
||||
from users.permissions import PermissionTag, require_permission
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
|
||||
class TalentRunStart(BaseModel):
|
||||
max_results: int | None = None
|
||||
location: str | None = None
|
||||
keywords: str | None = None
|
||||
|
||||
|
||||
@router.post("/talent/runs/start")
|
||||
async def start_talent_run(
|
||||
payload: TalentRunStart,
|
||||
job_post_id: str = Query(...),
|
||||
current_user: dict = Depends(require_permission(PermissionTag.TALENT_CREATE)),
|
||||
session: AsyncSession = Depends(get_session),
|
||||
):
|
||||
try:
|
||||
service = Talent(session=session)
|
||||
data = await service.start_run(
|
||||
job_post_id, payload.model_dump(exclude_unset=True), current_user
|
||||
)
|
||||
return JSONResponse(content={"data": data, "status_code": 200})
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
raise HTTPException(status_code=500, detail=str(e))
|
||||
|
||||
|
||||
@router.get("/talent/runs/status")
|
||||
async def talent_run_status(
|
||||
run_id: str = Query(...),
|
||||
current_user: dict = Depends(require_permission(PermissionTag.TALENT_VIEW)),
|
||||
session: AsyncSession = Depends(get_session),
|
||||
):
|
||||
try:
|
||||
service = Talent(session=session)
|
||||
data = await service.run_status(run_id)
|
||||
return JSONResponse(content={"data": data, "status_code": 200})
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
raise HTTPException(status_code=500, detail=str(e))
|
||||
|
||||
|
||||
@router.get("/talent/runs/fetch")
|
||||
async def fetch_talent_runs(
|
||||
job_post_id: str = Query(...),
|
||||
current_user: dict = Depends(require_permission(PermissionTag.TALENT_VIEW)),
|
||||
session: AsyncSession = Depends(get_session),
|
||||
):
|
||||
try:
|
||||
service = Talent(session=session)
|
||||
data, total = await service.fetch_runs(job_post_id)
|
||||
return JSONResponse(content={"data": data, "total": total, "status_code": 200})
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
raise HTTPException(status_code=500, detail=str(e))
|
||||
|
||||
|
||||
@router.get("/talent/profiles/fetch")
|
||||
async def fetch_talent_profiles(
|
||||
job_post_id: str = Query(...),
|
||||
search: str | None = Query(None),
|
||||
top: int | None = Query(None, ge=1, le=500),
|
||||
skip: int = Query(0, ge=0),
|
||||
current_user: dict = Depends(require_permission(PermissionTag.TALENT_VIEW)),
|
||||
session: AsyncSession = Depends(get_session),
|
||||
):
|
||||
try:
|
||||
service = Talent(session=session)
|
||||
data, total = await service.fetch_profiles(job_post_id, search=search, top=top, skip=skip)
|
||||
return JSONResponse(content={"data": data, "total": total, "status_code": 200})
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
raise HTTPException(status_code=500, detail=str(e))
|
||||
|
||||
|
||||
@router.get("/talent/profiles/fetch_by_id")
|
||||
async def fetch_talent_profile(
|
||||
profile_id: str = Query(...),
|
||||
current_user: dict = Depends(require_permission(PermissionTag.TALENT_VIEW)),
|
||||
session: AsyncSession = Depends(get_session),
|
||||
):
|
||||
try:
|
||||
service = Talent(session=session)
|
||||
data = await service.get_profile(profile_id)
|
||||
return JSONResponse(content={"data": data, "status_code": 200})
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
raise HTTPException(status_code=500, detail=str(e))
|
||||
|
||||
|
||||
@router.delete("/talent/profiles/delete")
|
||||
async def delete_talent_profile(
|
||||
profile_id: str = Query(...),
|
||||
current_user: dict = Depends(require_permission(PermissionTag.TALENT_DELETE)),
|
||||
session: AsyncSession = Depends(get_session),
|
||||
):
|
||||
try:
|
||||
service = Talent(session=session)
|
||||
data = await service.delete_profile(profile_id)
|
||||
return JSONResponse(content={"data": data, "status_code": 200})
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
raise HTTPException(status_code=500, detail=str(e))
|
||||
|
|
@ -0,0 +1,298 @@
|
|||
"""Talent sourcing tables: Apify actor runs and the LinkedIn profiles they find.
|
||||
|
||||
`talent_runs` is one row per paid actor run (vendor-id trio mirrors the Buffer
|
||||
columns on job_posts). `talent_profiles` is deduped per job by normalized
|
||||
LinkedIn URL across re-runs; `raw` keeps the full dataset item verbatim because
|
||||
actor output fields vary between actors and versions.
|
||||
"""
|
||||
|
||||
import uuid
|
||||
from datetime import datetime, timezone
|
||||
|
||||
from sqlalchemy import DateTime, JSON, UniqueConstraint, func
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
from sqlmodel import Field, SQLModel, select
|
||||
|
||||
|
||||
def _now() -> datetime:
|
||||
return datetime.now(timezone.utc)
|
||||
|
||||
|
||||
# Local run lifecycle. `pending` exists only between row insert and the Apify
|
||||
# start call succeeding; everything after start is driven by Apify's status.
|
||||
TERMINAL_RUN_STATUSES = ("succeeded", "failed", "timed_out", "aborted")
|
||||
|
||||
# Profile fields refreshed when a later run re-finds the same person. Kept at
|
||||
# module level: an underscore-prefixed class attribute on a SQLModel becomes a
|
||||
# Pydantic ModelPrivateAttr, which is not iterable. `is_deleted` is deliberately
|
||||
# absent — a dismissed profile stays dismissed.
|
||||
MUTABLE_PROFILE_FIELDS = (
|
||||
"public_id", "full_name", "headline", "location",
|
||||
"current_title", "current_company", "avatar_url", "summary", "skills",
|
||||
"match_score", "raw",
|
||||
)
|
||||
|
||||
|
||||
class TalentRuns(SQLModel, table=True):
|
||||
__tablename__ = "talent_runs"
|
||||
|
||||
id: uuid.UUID = Field(default_factory=uuid.uuid4, primary_key=True)
|
||||
job_post_id: uuid.UUID = Field(index=True, foreign_key="job_posts.id")
|
||||
requested_by: uuid.UUID | None = Field(default=None, foreign_key="users.id")
|
||||
status: str = Field(default="pending")
|
||||
actor_id: str = Field(default="")
|
||||
search_input: dict = Field(default_factory=dict, sa_type=JSON)
|
||||
max_results: int = Field(default=0)
|
||||
apify_run_id: str | None = Field(default=None)
|
||||
apify_dataset_id: str | None = Field(default=None)
|
||||
apify_error: str | None = Field(default=None)
|
||||
profiles_found: int = Field(default=0)
|
||||
started_at: datetime | None = Field(default=None, sa_type=DateTime(timezone=True))
|
||||
finished_at: datetime | None = Field(default=None, sa_type=DateTime(timezone=True))
|
||||
created_at: datetime = Field(default_factory=_now, sa_type=DateTime(timezone=True))
|
||||
updated_at: datetime = Field(default_factory=_now, sa_type=DateTime(timezone=True))
|
||||
is_deleted: bool = Field(default=False)
|
||||
|
||||
@staticmethod
|
||||
def _as_uuid(record_id) -> uuid.UUID | None:
|
||||
if record_id in (None, ""):
|
||||
return None
|
||||
try:
|
||||
return uuid.UUID(str(record_id))
|
||||
except ValueError:
|
||||
return None
|
||||
|
||||
@classmethod
|
||||
async def get_by_id(cls, session: AsyncSession, record_id):
|
||||
uid = cls._as_uuid(record_id)
|
||||
if uid is None:
|
||||
return None
|
||||
statement = select(cls).where(cls.id == uid, cls.is_deleted == False) # noqa: E712
|
||||
result = await session.execute(statement)
|
||||
return result.scalars().first()
|
||||
|
||||
@classmethod
|
||||
async def fetch_runs(cls, session: AsyncSession, *, job_post_id):
|
||||
jid = cls._as_uuid(job_post_id)
|
||||
if jid is None:
|
||||
return [], 0
|
||||
statement = select(cls).where(
|
||||
cls.job_post_id == jid, cls.is_deleted == False # noqa: E712
|
||||
)
|
||||
count_statement = select(func.count()).select_from(statement.subquery())
|
||||
total = (await session.execute(count_statement)).scalar_one()
|
||||
statement = statement.order_by(cls.created_at.desc())
|
||||
result = await session.execute(statement)
|
||||
return list(result.scalars().all()), total
|
||||
|
||||
@classmethod
|
||||
async def latest_active_run(cls, session: AsyncSession, job_post_id):
|
||||
jid = cls._as_uuid(job_post_id)
|
||||
if jid is None:
|
||||
return None
|
||||
statement = (
|
||||
select(cls)
|
||||
.where(
|
||||
cls.job_post_id == jid,
|
||||
cls.is_deleted == False, # noqa: E712
|
||||
cls.status.not_in(TERMINAL_RUN_STATUSES),
|
||||
)
|
||||
.order_by(cls.created_at.desc())
|
||||
)
|
||||
result = await session.execute(statement)
|
||||
return result.scalars().first()
|
||||
|
||||
@classmethod
|
||||
async def insert_run(cls, session: AsyncSession, fields: dict):
|
||||
row = cls(**fields)
|
||||
session.add(row)
|
||||
await session.commit()
|
||||
return await cls.get_by_id(session, row.id)
|
||||
|
||||
@classmethod
|
||||
async def _update(cls, session: AsyncSession, record_id, fields: dict):
|
||||
row = await cls.get_by_id(session, record_id)
|
||||
if not row:
|
||||
return None
|
||||
for key, value in fields.items():
|
||||
setattr(row, key, value)
|
||||
row.updated_at = _now()
|
||||
session.add(row)
|
||||
await session.commit()
|
||||
await session.refresh(row)
|
||||
return row
|
||||
|
||||
@classmethod
|
||||
async def mark_started(cls, session: AsyncSession, record_id, *, apify_run_id, apify_dataset_id):
|
||||
return await cls._update(session, record_id, {
|
||||
"status": "running",
|
||||
"apify_run_id": apify_run_id,
|
||||
"apify_dataset_id": apify_dataset_id,
|
||||
"started_at": _now(),
|
||||
})
|
||||
|
||||
@classmethod
|
||||
async def mark_rearmed(
|
||||
cls, session: AsyncSession, record_id, *,
|
||||
apify_run_id, apify_dataset_id, search_input: dict, found_so_far: int,
|
||||
):
|
||||
"""Point the SAME run row at a broadened follow-up actor run.
|
||||
|
||||
Status stays "running" so the frontend keeps polling and the 409
|
||||
active-run guard keeps holding; profiles_found accumulates across
|
||||
the ladder's batches.
|
||||
"""
|
||||
return await cls._update(session, record_id, {
|
||||
"status": "running",
|
||||
"apify_run_id": apify_run_id,
|
||||
"apify_dataset_id": apify_dataset_id,
|
||||
"search_input": search_input,
|
||||
"profiles_found": found_so_far,
|
||||
})
|
||||
|
||||
@classmethod
|
||||
async def mark_status(cls, session: AsyncSession, record_id, status: str):
|
||||
fields: dict = {"status": status}
|
||||
if status in TERMINAL_RUN_STATUSES:
|
||||
fields["finished_at"] = _now()
|
||||
return await cls._update(session, record_id, fields)
|
||||
|
||||
@classmethod
|
||||
async def mark_failed(cls, session: AsyncSession, record_id, error: str, *, status: str = "failed"):
|
||||
return await cls._update(session, record_id, {
|
||||
"status": status,
|
||||
"apify_error": (error or "")[:2000],
|
||||
"finished_at": _now(),
|
||||
})
|
||||
|
||||
@classmethod
|
||||
async def mark_succeeded(cls, session: AsyncSession, record_id, *, profiles_found: int):
|
||||
return await cls._update(session, record_id, {
|
||||
"status": "succeeded",
|
||||
"profiles_found": profiles_found,
|
||||
"apify_error": None,
|
||||
"finished_at": _now(),
|
||||
})
|
||||
|
||||
|
||||
class TalentProfiles(SQLModel, table=True):
|
||||
__tablename__ = "talent_profiles"
|
||||
__table_args__ = (
|
||||
UniqueConstraint("job_post_id", "linkedin_url", name="uq_talent_profiles_job_url"),
|
||||
)
|
||||
|
||||
id: uuid.UUID = Field(default_factory=uuid.uuid4, primary_key=True)
|
||||
job_post_id: uuid.UUID = Field(index=True, foreign_key="job_posts.id")
|
||||
run_id: uuid.UUID = Field(foreign_key="talent_runs.id")
|
||||
last_run_id: uuid.UUID | None = Field(default=None)
|
||||
linkedin_url: str
|
||||
public_id: str | None = Field(default=None)
|
||||
full_name: str | None = Field(default=None)
|
||||
headline: str | None = Field(default=None)
|
||||
location: str | None = Field(default=None)
|
||||
current_title: str | None = Field(default=None)
|
||||
current_company: str | None = Field(default=None)
|
||||
avatar_url: str | None = Field(default=None)
|
||||
summary: str | None = Field(default=None)
|
||||
skills: list = Field(default_factory=list, sa_type=JSON)
|
||||
match_score: int | None = Field(default=None)
|
||||
raw: dict = Field(default_factory=dict, sa_type=JSON)
|
||||
first_seen_at: datetime = Field(default_factory=_now, sa_type=DateTime(timezone=True))
|
||||
last_seen_at: datetime = Field(default_factory=_now, sa_type=DateTime(timezone=True))
|
||||
created_at: datetime = Field(default_factory=_now, sa_type=DateTime(timezone=True))
|
||||
updated_at: datetime = Field(default_factory=_now, sa_type=DateTime(timezone=True))
|
||||
is_deleted: bool = Field(default=False)
|
||||
|
||||
@classmethod
|
||||
async def fetch_profiles(
|
||||
cls, session: AsyncSession, *, job_post_id, search=None, top=None, skip=0
|
||||
):
|
||||
jid = TalentRuns._as_uuid(job_post_id)
|
||||
if jid is None:
|
||||
return [], 0
|
||||
statement = select(cls).where(
|
||||
cls.job_post_id == jid, cls.is_deleted == False # noqa: E712
|
||||
)
|
||||
if search:
|
||||
pattern = f"%{search}%"
|
||||
statement = statement.where(
|
||||
cls.full_name.ilike(pattern)
|
||||
| cls.headline.ilike(pattern)
|
||||
| cls.current_company.ilike(pattern)
|
||||
)
|
||||
count_statement = select(func.count()).select_from(statement.subquery())
|
||||
total = (await session.execute(count_statement)).scalar_one()
|
||||
statement = statement.order_by(
|
||||
cls.match_score.desc().nulls_last(),
|
||||
cls.last_seen_at.desc(),
|
||||
cls.created_at.desc(),
|
||||
)
|
||||
if skip:
|
||||
statement = statement.offset(skip)
|
||||
if top:
|
||||
statement = statement.limit(top)
|
||||
result = await session.execute(statement)
|
||||
return list(result.scalars().all()), total
|
||||
|
||||
@classmethod
|
||||
async def upsert_from_items(
|
||||
cls, session: AsyncSession, *, job_post_id, run_id, normalized_items: list[dict]
|
||||
) -> int:
|
||||
"""Insert new profiles, refresh re-found ones. One commit for the batch.
|
||||
|
||||
Dedupe key is (job_post_id, linkedin_url); dismissed rows are refreshed
|
||||
too but keep is_deleted=True so a re-run cannot resurrect them.
|
||||
"""
|
||||
jid = TalentRuns._as_uuid(job_post_id)
|
||||
rid = TalentRuns._as_uuid(run_id)
|
||||
persisted = 0
|
||||
for item in normalized_items:
|
||||
url = item.get("linkedin_url")
|
||||
if not url:
|
||||
continue
|
||||
statement = select(cls).where(
|
||||
cls.job_post_id == jid, cls.linkedin_url == url
|
||||
)
|
||||
existing = (await session.execute(statement)).scalars().first()
|
||||
if existing:
|
||||
for key in MUTABLE_PROFILE_FIELDS:
|
||||
if item.get(key) is not None:
|
||||
setattr(existing, key, item[key])
|
||||
existing.last_run_id = rid
|
||||
existing.last_seen_at = _now()
|
||||
existing.updated_at = _now()
|
||||
session.add(existing)
|
||||
else:
|
||||
session.add(cls(
|
||||
job_post_id=jid,
|
||||
run_id=rid,
|
||||
last_run_id=rid,
|
||||
linkedin_url=url,
|
||||
**{key: item.get(key) for key in MUTABLE_PROFILE_FIELDS},
|
||||
))
|
||||
persisted += 1
|
||||
await session.commit()
|
||||
return persisted
|
||||
|
||||
@classmethod
|
||||
async def get_profile_by_id(cls, session: AsyncSession, record_id):
|
||||
uid = TalentRuns._as_uuid(record_id)
|
||||
if uid is None:
|
||||
return None
|
||||
statement = select(cls).where(cls.id == uid, cls.is_deleted == False) # noqa: E712
|
||||
return (await session.execute(statement)).scalars().first()
|
||||
|
||||
@classmethod
|
||||
async def soft_delete_profile(cls, session: AsyncSession, record_id):
|
||||
row = await cls.get_profile_by_id(session, record_id)
|
||||
if not row:
|
||||
return None
|
||||
row.is_deleted = True
|
||||
row.updated_at = _now()
|
||||
session.add(row)
|
||||
await session.commit()
|
||||
return row
|
||||
|
||||
|
||||
import job.job_post.models as _job_post_models # noqa: E402, F401
|
||||
import users.models as _users_models # noqa: E402, F401
|
||||
|
|
@ -0,0 +1,554 @@
|
|||
"""Apify REST helpers and LinkedIn profile normalization.
|
||||
|
||||
Pure module: no FastAPI imports and no HTTPException.
|
||||
|
||||
The default actor is HarvestAPI's no-cookie LinkedIn people search
|
||||
(harvestapi~linkedin-profile-search). Its input schema was verified live:
|
||||
`searchQuery` (fuzzy string), `maxItems` (int), `locations` (array of strings),
|
||||
`profileScraperMode` ("Short" | "Full" | "Full + email search"). Swapping actors
|
||||
later means changing APIFY_ACTOR_ID plus, at most, build_actor_input and
|
||||
normalize_profile.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import re
|
||||
from urllib.parse import urlsplit
|
||||
|
||||
import httpx
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
|
||||
# The user's console .env entry is APIFY_TOKEN; APIFY_API_TOKEN is the documented name.
|
||||
APIFY_API_TOKEN = os.getenv("APIFY_API_TOKEN") or os.getenv("APIFY_TOKEN")
|
||||
APIFY_API_BASE = os.getenv("APIFY_API_BASE", "https://api.apify.com/v2")
|
||||
APIFY_ACTOR_ID = os.getenv("APIFY_ACTOR_ID", "harvestapi~linkedin-profile-search")
|
||||
APIFY_MAX_RESULTS = int(os.getenv("APIFY_MAX_RESULTS", "25"))
|
||||
APIFY_PROFILE_MODE = os.getenv("APIFY_PROFILE_MODE", "Full")
|
||||
APIFY_TIMEOUT = float(os.getenv("APIFY_TIMEOUT", "30"))
|
||||
# Server-side spend ceiling per run (maxTotalChargeUsd). The HarvestAPI actor is
|
||||
# pay-per-EVENT ($0.10/search page + per-profile), so Apify's maxItems billing
|
||||
# param does not apply — it was rejected live with "Maximum cost per run is less
|
||||
# than the allowed minimum of $0.10". A 25-profile Full run costs ~$0.20.
|
||||
APIFY_MAX_COST_USD = float(os.getenv("APIFY_MAX_COST_USD", "1.0"))
|
||||
|
||||
|
||||
def _csv_env(name: str, default: str) -> list[str]:
|
||||
return [s.strip() for s in os.getenv(name, default).split(",") if s.strip()]
|
||||
|
||||
|
||||
# The user's own companies: their CURRENT employees must never appear in sourced
|
||||
# results. Names drive the always-on server-side filter (case-insensitive
|
||||
# substring, so "Utopia Brands Pakistan" matches too). URLs drive the actor's
|
||||
# excludeCurrentCompanies filter, which wants full LinkedIn company URLs and
|
||||
# stops those profiles from being scraped (and paid for) at all.
|
||||
APIFY_EXCLUDE_COMPANIES = _csv_env("APIFY_EXCLUDE_COMPANIES", "Utopia Brands,Utopia Deals")
|
||||
APIFY_EXCLUDE_COMPANY_URLS = _csv_env(
|
||||
"APIFY_EXCLUDE_COMPANY_URLS",
|
||||
"https://www.linkedin.com/company/utopiadeals,"
|
||||
"https://www.linkedin.com/company/utopia-brands-usa,"
|
||||
"https://www.linkedin.com/company/utopiabrands",
|
||||
)
|
||||
|
||||
|
||||
def _matches_excluded(text) -> bool:
|
||||
haystack = " ".join(str(text or "").lower().split())
|
||||
return bool(haystack) and any(
|
||||
name.lower() in haystack for name in APIFY_EXCLUDE_COMPANIES
|
||||
)
|
||||
|
||||
|
||||
def is_excluded_profile(profile: dict) -> bool:
|
||||
"""True when the person currently works at one of the excluded companies.
|
||||
|
||||
The headline is only consulted when no current company was extracted, so an
|
||||
"ex-Utopia" headline on someone now elsewhere does not exclude them.
|
||||
"""
|
||||
company = (profile or {}).get("current_company")
|
||||
if _matches_excluded(company):
|
||||
return True
|
||||
return not company and _matches_excluded((profile or {}).get("headline"))
|
||||
|
||||
# Apify run status -> talent_runs.status. Transitional states stay "running";
|
||||
# unknown values also stay "running" so we never commit a terminal state we
|
||||
# don't understand (Buffer precedent).
|
||||
APIFY_STATUS_TO_LOCAL = {
|
||||
"READY": "running",
|
||||
"RUNNING": "running",
|
||||
"TIMING-OUT": "running",
|
||||
"ABORTING": "running",
|
||||
"SUCCEEDED": "succeeded",
|
||||
"FAILED": "failed",
|
||||
"TIMED-OUT": "timed_out",
|
||||
"ABORTED": "aborted",
|
||||
}
|
||||
|
||||
TERMINAL_STATUSES = {"succeeded", "failed", "timed_out", "aborted"}
|
||||
|
||||
|
||||
def local_status(apify_status) -> str:
|
||||
return APIFY_STATUS_TO_LOCAL.get(str(apify_status or "").upper(), "running")
|
||||
|
||||
|
||||
class ApifyError(RuntimeError):
|
||||
def __init__(self, message: str, *, code: str | None = None):
|
||||
super().__init__(message)
|
||||
self.code = code
|
||||
|
||||
|
||||
def _headers() -> dict:
|
||||
if not APIFY_API_TOKEN:
|
||||
raise RuntimeError("APIFY_API_TOKEN is not configured")
|
||||
return {
|
||||
"Authorization": f"Bearer {APIFY_API_TOKEN}",
|
||||
"Content-Type": "application/json",
|
||||
}
|
||||
|
||||
|
||||
def _raise_for_response(response: httpx.Response) -> None:
|
||||
if response.status_code < 400:
|
||||
return
|
||||
try:
|
||||
error = (response.json() or {}).get("error") or {}
|
||||
except ValueError:
|
||||
error = {}
|
||||
if error.get("message"):
|
||||
raise ApifyError(error["message"], code=error.get("type"))
|
||||
raise httpx.HTTPStatusError(
|
||||
response.text, request=response.request, response=response
|
||||
)
|
||||
|
||||
|
||||
# LinkedIn's years-of-experience facet, as the actor's yearsOfExperienceIds
|
||||
# enum defines it (verified from the actor's input schema): id -> (min, max)
|
||||
# whole years. A job's experience_min/experience_max selects every overlapping
|
||||
# bucket.
|
||||
EXPERIENCE_BUCKETS = {
|
||||
"1": (0, 0), # Less than 1 year
|
||||
"2": (1, 2), # 1 to 2 years
|
||||
"3": (3, 5), # 3 to 5 years
|
||||
"4": (6, 10), # 6 to 10 years
|
||||
"5": (11, 60), # More than 10 years
|
||||
}
|
||||
|
||||
|
||||
def years_of_experience_ids(experience_min, experience_max) -> list[str]:
|
||||
"""Bucket ids overlapping [experience_min, experience_max]; [] = no filter."""
|
||||
if experience_min is None and experience_max is None:
|
||||
return []
|
||||
lo = int(experience_min) if experience_min is not None else 0
|
||||
hi = int(experience_max) if experience_max is not None else 60
|
||||
if hi < lo:
|
||||
lo, hi = hi, lo
|
||||
return [
|
||||
bucket_id
|
||||
for bucket_id, (b_lo, b_hi) in EXPERIENCE_BUCKETS.items()
|
||||
if b_hi >= lo and b_lo <= hi
|
||||
]
|
||||
|
||||
|
||||
# Job "locations" that are work arrangements, not places. Sending one as the
|
||||
# actor's locations filter returns an empty dataset — verified live: a run with
|
||||
# locations=["Remote"] found 0 profiles where the same query with a real
|
||||
# geography found 5. Filter them out instead of filtering by them.
|
||||
NON_GEOGRAPHIC_LOCATIONS = {
|
||||
"remote", "hybrid", "onsite", "on-site", "on site",
|
||||
"anywhere", "flexible", "wfh", "work from home",
|
||||
}
|
||||
|
||||
|
||||
def _geographic_location(value) -> str | None:
|
||||
text = str(value or "").strip()
|
||||
if not text or text.lower() in NON_GEOGRAPHIC_LOCATIONS:
|
||||
return None
|
||||
return text
|
||||
|
||||
|
||||
def _skill_terms(*entry_lists) -> list[str]:
|
||||
"""Keyword-like entries only (max 3 words, 30 chars), deduped in order.
|
||||
|
||||
Job requirements are sometimes skills ("Python", "Amazon Seller Central")
|
||||
and sometimes prose ("2-5 years of experience managing Amazon PPC
|
||||
campaigns..."). Prose in the fuzzy searchQuery strangles it — verified
|
||||
live: a sentence-stuffed query matched 2 people country-wide and 0 in
|
||||
Karachi, where the title alone finds plenty.
|
||||
"""
|
||||
terms: list[str] = []
|
||||
seen: set[str] = set()
|
||||
for entries in entry_lists:
|
||||
for entry in entries or []:
|
||||
text = " ".join(str(entry).split())
|
||||
if not text or text.lower() in seen:
|
||||
continue
|
||||
if len(text) <= 30 and len(text.split()) <= 3:
|
||||
terms.append(text)
|
||||
seen.add(text.lower())
|
||||
return terms
|
||||
|
||||
|
||||
def build_actor_input(
|
||||
job: dict, *, max_results: int, overrides: dict | None = None, start_page: int = 1
|
||||
) -> dict:
|
||||
"""Deterministic actor input from job fields. No LLM involved.
|
||||
|
||||
The job title goes into LinkedIn's CURRENT-TITLE facet (currentJobTitles),
|
||||
not the keyword box: a keyword query matches words anywhere in a profile,
|
||||
so "AI Engineer Python..." returned a pool that was 52% generic software
|
||||
engineers (every full-stack profile mentions Python). Verified live: the
|
||||
facet alone returns full pages of genuinely AI-titled people. The keyword
|
||||
box carries only the skill terms. start_page > 1 continues a previous
|
||||
search deeper into the result pages (25 profiles per page), so a re-run
|
||||
surfaces new people instead of re-finding the first page.
|
||||
"""
|
||||
overrides = overrides or {}
|
||||
title = str(job.get("title") or "").strip()[:100]
|
||||
terms = _skill_terms(job.get("requirements"), job.get("optional_skills"))[:3]
|
||||
query = " ".join(terms).strip()[:200]
|
||||
if overrides.get("keywords"):
|
||||
query = str(overrides["keywords"]).strip()[:200]
|
||||
|
||||
actor_input: dict = {
|
||||
"maxItems": max_results,
|
||||
"profileScraperMode": APIFY_PROFILE_MODE,
|
||||
}
|
||||
if title:
|
||||
actor_input["currentJobTitles"] = [title]
|
||||
if query:
|
||||
actor_input["searchQuery"] = query
|
||||
elif not title:
|
||||
# No facet and no terms: nothing left to search by.
|
||||
actor_input["searchQuery"] = ""
|
||||
experience_ids = years_of_experience_ids(
|
||||
job.get("experience_min"), job.get("experience_max")
|
||||
)
|
||||
if experience_ids:
|
||||
actor_input["yearsOfExperienceIds"] = experience_ids
|
||||
if APIFY_EXCLUDE_COMPANY_URLS:
|
||||
actor_input["excludeCurrentCompanies"] = APIFY_EXCLUDE_COMPANY_URLS
|
||||
if start_page and int(start_page) > 1:
|
||||
actor_input["startPage"] = min(int(start_page), 100)
|
||||
if "location" in overrides and overrides["location"] is not None:
|
||||
# An explicit override wins outright — "Remote" here means the caller
|
||||
# wants no geography constraint, not a fallback to the job's location.
|
||||
location = _geographic_location(overrides["location"])
|
||||
else:
|
||||
location = _geographic_location(job.get("location"))
|
||||
if location:
|
||||
actor_input["locations"] = [location]
|
||||
return actor_input
|
||||
|
||||
|
||||
def broaden_actor_input(actor_input: dict) -> dict | None:
|
||||
"""Next rung of the thin-results broadening ladder, or None when exhausted.
|
||||
|
||||
A search ANDs facet + keywords + location + experience; in a single city
|
||||
that intersection can collapse to one person (seen live: Amazon PPC +
|
||||
Karachi returned 1). Rungs: (1) drop the keyword query, keeping the title
|
||||
facet; (2) drop the facet and search the title as keywords instead.
|
||||
Location, experience and company exclusions are never relaxed — they are
|
||||
user intent, not tuning.
|
||||
"""
|
||||
current = dict(actor_input)
|
||||
if current.get("currentJobTitles") and "searchQuery" in current:
|
||||
current.pop("searchQuery")
|
||||
return current
|
||||
if current.get("currentJobTitles"):
|
||||
title = current.pop("currentJobTitles")[0]
|
||||
current["searchQuery"] = title
|
||||
return current
|
||||
return None
|
||||
|
||||
|
||||
async def start_actor_run(actor_input: dict, *, actor_id: str | None = None) -> dict:
|
||||
"""POST /acts/{id}/runs. maxTotalChargeUsd caps spend on Apify's side,
|
||||
independent of what the actor does with its input; the profile count itself
|
||||
is limited by the maxItems field inside the actor input."""
|
||||
actor = actor_id or APIFY_ACTOR_ID
|
||||
async with httpx.AsyncClient(timeout=APIFY_TIMEOUT) as client:
|
||||
response = await client.post(
|
||||
f"{APIFY_API_BASE}/acts/{actor}/runs",
|
||||
params={"maxTotalChargeUsd": APIFY_MAX_COST_USD},
|
||||
json=actor_input,
|
||||
headers=_headers(),
|
||||
)
|
||||
_raise_for_response(response)
|
||||
data = (response.json() or {}).get("data") or {}
|
||||
if not data.get("id"):
|
||||
raise ApifyError("Apify did not return a run id")
|
||||
return data
|
||||
|
||||
|
||||
async def get_run(run_id: str) -> dict:
|
||||
async with httpx.AsyncClient(timeout=APIFY_TIMEOUT) as client:
|
||||
response = await client.get(
|
||||
f"{APIFY_API_BASE}/actor-runs/{run_id}", headers=_headers()
|
||||
)
|
||||
_raise_for_response(response)
|
||||
return (response.json() or {}).get("data") or {}
|
||||
|
||||
|
||||
async def get_dataset_items(dataset_id: str, *, limit: int, offset: int = 0) -> list[dict]:
|
||||
async with httpx.AsyncClient(timeout=APIFY_TIMEOUT) as client:
|
||||
response = await client.get(
|
||||
f"{APIFY_API_BASE}/datasets/{dataset_id}/items",
|
||||
params={"format": "json", "clean": "true", "limit": limit, "offset": offset},
|
||||
headers=_headers(),
|
||||
)
|
||||
_raise_for_response(response)
|
||||
body = response.json()
|
||||
return body if isinstance(body, list) else []
|
||||
|
||||
|
||||
async def get_me() -> dict:
|
||||
"""Cheap auth sanity check; used by verification, not the request path."""
|
||||
async with httpx.AsyncClient(timeout=APIFY_TIMEOUT) as client:
|
||||
response = await client.get(f"{APIFY_API_BASE}/users/me", headers=_headers())
|
||||
_raise_for_response(response)
|
||||
return (response.json() or {}).get("data") or {}
|
||||
|
||||
|
||||
def normalize_linkedin_url(url) -> str | None:
|
||||
"""Canonical dedupe key: https, lowercase host/path, no query or trailing slash."""
|
||||
text = str(url or "").strip()
|
||||
if not text:
|
||||
return None
|
||||
if "//" not in text:
|
||||
text = f"https://{text}"
|
||||
parts = urlsplit(text)
|
||||
host = parts.netloc.lower()
|
||||
if "linkedin.com" not in host:
|
||||
return None
|
||||
path = parts.path.rstrip("/")
|
||||
return f"https://{host}{path}".lower()
|
||||
|
||||
|
||||
def _first_string(item: dict, *keys) -> str | None:
|
||||
for key in keys:
|
||||
value = item.get(key)
|
||||
if isinstance(value, str) and value.strip():
|
||||
return value.strip()
|
||||
return None
|
||||
|
||||
|
||||
def _location_text(value) -> str | None:
|
||||
if isinstance(value, str) and value.strip():
|
||||
return value.strip()
|
||||
if isinstance(value, dict):
|
||||
for key in ("linkedinText", "text", "name", "default"):
|
||||
nested = value.get(key)
|
||||
if isinstance(nested, str) and nested.strip():
|
||||
return nested.strip()
|
||||
return None
|
||||
|
||||
|
||||
def _photo_url(item: dict) -> str | None:
|
||||
for key in ("photo", "profilePicture", "avatar", "photoUrl", "profilePic", "image"):
|
||||
value = item.get(key)
|
||||
if isinstance(value, str) and value.strip():
|
||||
return value.strip()
|
||||
if isinstance(value, dict):
|
||||
nested = value.get("url")
|
||||
if isinstance(nested, str) and nested.strip():
|
||||
return nested.strip()
|
||||
return None
|
||||
|
||||
|
||||
def _skills(item: dict) -> list[str]:
|
||||
"""Up to 10 skill names; entries arrive as strings or {name: ...} dicts."""
|
||||
names: list[str] = []
|
||||
for key in ("topSkills", "skills"):
|
||||
for entry in item.get(key) or []:
|
||||
name = entry if isinstance(entry, str) else (
|
||||
entry.get("name") if isinstance(entry, dict) else None
|
||||
)
|
||||
if name and str(name).strip() and str(name).strip() not in names:
|
||||
names.append(str(name).strip())
|
||||
if names:
|
||||
break
|
||||
return names[:10]
|
||||
|
||||
|
||||
def _current_position(item: dict) -> tuple[str | None, str | None]:
|
||||
"""(title, company) from the most recent experience entry, however spelled."""
|
||||
position = item.get("position") or item.get("currentPosition")
|
||||
if isinstance(position, dict):
|
||||
title = _first_string(position, "title", "role")
|
||||
company = _first_string(position, "companyName", "company")
|
||||
if title or company:
|
||||
return title, company
|
||||
experience = item.get("experience") or item.get("experiences")
|
||||
if isinstance(experience, list) and experience:
|
||||
entry = experience[0]
|
||||
if isinstance(entry, dict):
|
||||
company = _first_string(entry, "companyName", "company")
|
||||
if company is None:
|
||||
nested = entry.get("company")
|
||||
if isinstance(nested, dict):
|
||||
company = _first_string(nested, "name")
|
||||
return _first_string(entry, "title", "position", "role"), company
|
||||
return None, _first_string(item, "companyName", "currentCompany")
|
||||
|
||||
|
||||
_TOKEN_STOPWORDS = {
|
||||
"and", "or", "the", "of", "for", "with", "in", "a", "an", "to",
|
||||
# Requirement-prose filler that appears in almost every profile and would
|
||||
# inflate every score equally, flattening the ranking.
|
||||
"experience", "years", "year", "strong", "including", "ability",
|
||||
"knowledge", "skills", "understanding", "familiarity", "proficiency",
|
||||
"hands", "must", "have", "plus", "good", "excellent", "etc",
|
||||
}
|
||||
|
||||
|
||||
def _clean_phrase(text) -> str:
|
||||
cleaned = re.sub(r"[^a-z0-9+#]+", " ", str(text or "").lower())
|
||||
return " ".join(
|
||||
t for t in cleaned.split() if len(t) > 1 and t not in _TOKEN_STOPWORDS
|
||||
)
|
||||
|
||||
|
||||
def _match_tokens(*texts) -> set[str]:
|
||||
tokens: set[str] = set()
|
||||
for text in texts:
|
||||
tokens.update(_clean_phrase(text).split())
|
||||
return tokens
|
||||
|
||||
|
||||
def relevance_score(job: dict, profile: dict) -> int:
|
||||
"""0-100 job-fit rank for sorting, computed when a profile is persisted.
|
||||
|
||||
Deterministic and free. Title component: the job title as an exact PHRASE
|
||||
in the person's current title scores 55, in their headline 45; scattered
|
||||
token overlap caps at 35 — a keyword-stuffed headline ("AI/ML Engineer |
|
||||
Python | FastAPI | ...") must not outrank someone whose title IS the job
|
||||
title, which is exactly what token overlap alone did on live data.
|
||||
|
||||
Skills component (up to 45): GRADED token overlap between the content
|
||||
words of the job's requirements + optional skills and the person's
|
||||
title/headline/skills/summary. Graded, not per-term all-or-nothing: the
|
||||
title facet makes every sourced profile earn the same title points, so
|
||||
all differentiation lives here — an all-or-nothing single term put a
|
||||
whole live pool on exactly 60.
|
||||
"""
|
||||
job_title = _clean_phrase(job.get("title"))
|
||||
title_text = _clean_phrase(profile.get("current_title"))
|
||||
headline_text = _clean_phrase(profile.get("headline"))
|
||||
if job_title and job_title in title_text:
|
||||
title_component = 55.0
|
||||
elif job_title and job_title in headline_text:
|
||||
title_component = 45.0
|
||||
else:
|
||||
title_tokens = set(job_title.split())
|
||||
role_tokens = set(title_text.split()) | set(headline_text.split())
|
||||
ratio = len(title_tokens & role_tokens) / len(title_tokens) if title_tokens else 0.0
|
||||
title_component = 35 * ratio
|
||||
|
||||
job_tokens = _match_tokens(
|
||||
*(job.get("requirements") or []), *(job.get("optional_skills") or [])
|
||||
)
|
||||
profile_tokens = _match_tokens(
|
||||
profile.get("current_title"),
|
||||
profile.get("headline"),
|
||||
" ".join(profile.get("skills") or []),
|
||||
profile.get("summary"),
|
||||
)
|
||||
skills_ratio = (
|
||||
len(job_tokens & profile_tokens) / len(job_tokens) if job_tokens else 0.0
|
||||
)
|
||||
|
||||
return round(title_component + 45 * skills_ratio)
|
||||
|
||||
|
||||
def _date_text(value) -> str | None:
|
||||
"""HarvestAPI dates arrive as {"month": "Jun", "year": 2025, "text": "Jun 2025"}."""
|
||||
if isinstance(value, dict):
|
||||
text = value.get("text")
|
||||
if isinstance(text, str) and text.strip():
|
||||
return text.strip()
|
||||
if isinstance(value, str) and value.strip():
|
||||
return value.strip()
|
||||
return None
|
||||
|
||||
|
||||
def _entry_company(entry: dict) -> str | None:
|
||||
company = _first_string(entry, "companyName")
|
||||
if company is None and isinstance(entry.get("company"), dict):
|
||||
company = _first_string(entry["company"], "name")
|
||||
return company
|
||||
|
||||
|
||||
def extract_experience(raw: dict) -> list[dict]:
|
||||
"""Employment history from a stored raw item, for the profile detail view."""
|
||||
entries: list[dict] = []
|
||||
for item in (raw or {}).get("experience") or []:
|
||||
if not isinstance(item, dict):
|
||||
continue
|
||||
start = _date_text(item.get("startDate"))
|
||||
end = _date_text(item.get("endDate"))
|
||||
description = str(item.get("description") or "").strip()
|
||||
entries.append({
|
||||
"title": _first_string(item, "position", "title", "role"),
|
||||
"company": _entry_company(item),
|
||||
"employment_type": _first_string(item, "employmentType"),
|
||||
"location": _location_text(item.get("location")),
|
||||
"duration": _first_string(item, "duration"),
|
||||
"period": " – ".join(p for p in (start, end) if p) or None,
|
||||
"description": description[:400] or None,
|
||||
"skills": _skills(item)[:6],
|
||||
})
|
||||
if len(entries) == 10:
|
||||
break
|
||||
return entries
|
||||
|
||||
|
||||
def extract_education(raw: dict) -> list[dict]:
|
||||
entries: list[dict] = []
|
||||
for item in (raw or {}).get("education") or []:
|
||||
if not isinstance(item, dict):
|
||||
continue
|
||||
start = _date_text(item.get("startDate"))
|
||||
end = _date_text(item.get("endDate"))
|
||||
entries.append({
|
||||
"school": _first_string(item, "schoolName", "school"),
|
||||
"degree": _first_string(item, "degree"),
|
||||
"field": _first_string(item, "fieldOfStudy", "field"),
|
||||
"period": _first_string(item, "period") or (" – ".join(p for p in (start, end) if p) or None),
|
||||
})
|
||||
if len(entries) == 5:
|
||||
break
|
||||
return entries
|
||||
|
||||
|
||||
def normalize_profile(item: dict) -> dict | None:
|
||||
"""Tolerant extraction of the card fields from one dataset item.
|
||||
|
||||
Returns None (skip, not fail) when the item has no LinkedIn URL. The full
|
||||
item always rides along as `raw` so nothing is lost to key drift.
|
||||
"""
|
||||
if not isinstance(item, dict):
|
||||
return None
|
||||
url = normalize_linkedin_url(
|
||||
_first_string(item, "linkedinUrl", "url", "profileUrl", "publicProfileUrl", "link")
|
||||
)
|
||||
if not url:
|
||||
return None
|
||||
name = _first_string(item, "fullName", "name")
|
||||
if not name:
|
||||
first = _first_string(item, "firstName") or ""
|
||||
last = _first_string(item, "lastName") or ""
|
||||
name = f"{first} {last}".strip() or None
|
||||
title, company = _current_position(item)
|
||||
return {
|
||||
"linkedin_url": url,
|
||||
"public_id": _first_string(item, "publicIdentifier", "publicId"),
|
||||
"full_name": name,
|
||||
"headline": _first_string(item, "headline", "subTitle", "occupation"),
|
||||
"location": _location_text(item.get("location")),
|
||||
"current_title": title,
|
||||
"current_company": company,
|
||||
"avatar_url": _photo_url(item),
|
||||
"summary": _first_string(item, "about", "summary"),
|
||||
"skills": _skills(item),
|
||||
"raw": item,
|
||||
}
|
||||
|
|
@ -0,0 +1,50 @@
|
|||
from talent.plugins import extract_education, extract_experience
|
||||
|
||||
|
||||
def serialize_talent_run(row) -> dict:
|
||||
return {
|
||||
"id": str(row.id) if row.id else None,
|
||||
"job_post_id": str(row.job_post_id) if row.job_post_id else None,
|
||||
"status": row.status,
|
||||
"actor_id": row.actor_id,
|
||||
"search_input": row.search_input or {},
|
||||
"max_results": row.max_results,
|
||||
"profiles_found": row.profiles_found,
|
||||
"apify_run_id": row.apify_run_id,
|
||||
"apify_error": row.apify_error,
|
||||
"started_at": row.started_at.isoformat() if row.started_at else None,
|
||||
"finished_at": row.finished_at.isoformat() if row.finished_at else None,
|
||||
"created_at": row.created_at.isoformat() if row.created_at else None,
|
||||
}
|
||||
|
||||
|
||||
def serialize_talent_profile(row) -> dict:
|
||||
# `raw` stays server-side: it is an actor-shaped blob that can be large and
|
||||
# is only needed for debugging/re-mapping, not for the profile cards.
|
||||
return {
|
||||
"id": str(row.id) if row.id else None,
|
||||
"job_post_id": str(row.job_post_id) if row.job_post_id else None,
|
||||
"linkedin_url": row.linkedin_url,
|
||||
"public_id": row.public_id,
|
||||
"full_name": row.full_name,
|
||||
"headline": row.headline,
|
||||
"location": row.location,
|
||||
"current_title": row.current_title,
|
||||
"current_company": row.current_company,
|
||||
"avatar_url": row.avatar_url,
|
||||
"summary": row.summary,
|
||||
"skills": row.skills or [],
|
||||
"match_score": row.match_score,
|
||||
"first_seen_at": row.first_seen_at.isoformat() if row.first_seen_at else None,
|
||||
"last_seen_at": row.last_seen_at.isoformat() if row.last_seen_at else None,
|
||||
}
|
||||
|
||||
|
||||
def serialize_talent_profile_detail(row) -> dict:
|
||||
# The card payload plus employment/education history unpacked from the raw
|
||||
# actor item. Detail is fetched one profile at a time, so the extra weight
|
||||
# never rides along with the list endpoint.
|
||||
data = serialize_talent_profile(row)
|
||||
data["experience"] = extract_experience(row.raw or {})
|
||||
data["education"] = extract_education(row.raw or {})
|
||||
return data
|
||||
|
|
@ -0,0 +1,220 @@
|
|||
import httpx
|
||||
from fastapi import HTTPException
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from job.job_post.models import JobPosts
|
||||
from talent import plugins
|
||||
from talent.models import TalentProfiles, TalentRuns
|
||||
from talent.serializers import (
|
||||
serialize_talent_profile,
|
||||
serialize_talent_profile_detail,
|
||||
serialize_talent_run,
|
||||
)
|
||||
|
||||
|
||||
def _search_basis(actor_input: dict) -> dict:
|
||||
"""The identity of a search, ignoring pagination and batch-size knobs."""
|
||||
return {k: v for k, v in (actor_input or {}).items() if k not in ("startPage", "maxItems")}
|
||||
|
||||
|
||||
class Talent:
|
||||
def __init__(self, session: AsyncSession):
|
||||
self.session = session
|
||||
|
||||
async def _get_job(self, job_post_id):
|
||||
job = await JobPosts.get_job_post_by_id(self.session, job_post_id)
|
||||
if not job or job.is_deleted:
|
||||
raise HTTPException(status_code=404, detail="Job post not found")
|
||||
return job
|
||||
|
||||
async def start_run(self, job_post_id, payload, current_user):
|
||||
job = await self._get_job(job_post_id)
|
||||
|
||||
active = await TalentRuns.latest_active_run(self.session, job_post_id)
|
||||
if active:
|
||||
raise HTTPException(
|
||||
status_code=409,
|
||||
detail="A talent search is already running for this job",
|
||||
)
|
||||
|
||||
requested = payload.get("max_results")
|
||||
max_results = min(int(requested), plugins.APIFY_MAX_RESULTS) if requested else plugins.APIFY_MAX_RESULTS
|
||||
if max_results < 1:
|
||||
raise HTTPException(status_code=422, detail="max_results must be at least 1")
|
||||
# Floor of 10 per paid run (user asked for at least 10 results a
|
||||
# search) — unless the env cap itself is set lower.
|
||||
max_results = max(max_results, min(10, plugins.APIFY_MAX_RESULTS))
|
||||
|
||||
overrides = {
|
||||
"keywords": payload.get("keywords"),
|
||||
"location": payload.get("location"),
|
||||
}
|
||||
job_fields = {
|
||||
"title": job.title,
|
||||
"requirements": job.requirements,
|
||||
"optional_skills": job.optional_skills,
|
||||
"location": job.location,
|
||||
"experience_min": job.experience_min,
|
||||
"experience_max": job.experience_max,
|
||||
}
|
||||
actor_input = plugins.build_actor_input(
|
||||
job_fields, max_results=max_results, overrides=overrides
|
||||
)
|
||||
|
||||
# Re-running the same search continues deeper into LinkedIn's result
|
||||
# pages (25 profiles each), so every run surfaces new people. A changed
|
||||
# query/location/experience is a different search and restarts at page 1.
|
||||
basis = _search_basis(actor_input)
|
||||
prior_runs, _ = await TalentRuns.fetch_runs(self.session, job_post_id=job_post_id)
|
||||
prior_pages = [
|
||||
int((r.search_input or {}).get("startPage") or 1)
|
||||
for r in prior_runs
|
||||
if r.status == "succeeded" and _search_basis(r.search_input) == basis
|
||||
]
|
||||
if prior_pages:
|
||||
actor_input = plugins.build_actor_input(
|
||||
job_fields,
|
||||
max_results=max_results,
|
||||
overrides=overrides,
|
||||
start_page=max(prior_pages) + 1,
|
||||
)
|
||||
run = await TalentRuns.insert_run(self.session, {
|
||||
"job_post_id": job.id,
|
||||
"requested_by": TalentRuns._as_uuid((current_user or {}).get("id")),
|
||||
"status": "pending",
|
||||
"actor_id": plugins.APIFY_ACTOR_ID,
|
||||
"search_input": actor_input,
|
||||
"max_results": max_results,
|
||||
})
|
||||
|
||||
try:
|
||||
started = await plugins.start_actor_run(actor_input)
|
||||
except (httpx.HTTPError, plugins.ApifyError, RuntimeError) as exc:
|
||||
# Keep the failed row for run history, then surface the vendor error.
|
||||
await TalentRuns.mark_failed(self.session, run.id, str(exc))
|
||||
raise HTTPException(status_code=502, detail=f"Apify run could not be started: {exc}")
|
||||
|
||||
run = await TalentRuns.mark_started(
|
||||
self.session,
|
||||
run.id,
|
||||
apify_run_id=started.get("id"),
|
||||
apify_dataset_id=started.get("defaultDatasetId"),
|
||||
)
|
||||
return serialize_talent_run(run)
|
||||
|
||||
async def run_status(self, run_id):
|
||||
run = await TalentRuns.get_by_id(self.session, run_id)
|
||||
if not run:
|
||||
raise HTTPException(status_code=404, detail="Talent run not found")
|
||||
|
||||
# Terminal runs are immutable: no Apify call, no re-persist. This makes
|
||||
# the poll endpoint idempotent and cheap once a run has settled.
|
||||
if run.status in plugins.TERMINAL_STATUSES:
|
||||
return serialize_talent_run(run)
|
||||
|
||||
if not run.apify_run_id:
|
||||
# pending row whose start call never completed (crash between insert
|
||||
# and mark_started): nothing to poll, mark it failed.
|
||||
run = await TalentRuns.mark_failed(
|
||||
self.session, run.id, "Run was never started on Apify"
|
||||
)
|
||||
return serialize_talent_run(run)
|
||||
|
||||
try:
|
||||
remote = await plugins.get_run(run.apify_run_id)
|
||||
except (httpx.HTTPError, plugins.ApifyError, RuntimeError) as exc:
|
||||
if isinstance(exc, plugins.ApifyError) and exc.code == "record-not-found":
|
||||
run = await TalentRuns.mark_failed(
|
||||
self.session, run.id, "Apify run no longer exists"
|
||||
)
|
||||
return serialize_talent_run(run)
|
||||
raise HTTPException(status_code=502, detail=f"Apify status check failed: {exc}")
|
||||
|
||||
status = plugins.local_status(remote.get("status"))
|
||||
if status == "running":
|
||||
run = await TalentRuns.mark_status(self.session, run.id, "running")
|
||||
return serialize_talent_run(run)
|
||||
|
||||
if status == "succeeded":
|
||||
dataset_id = run.apify_dataset_id or remote.get("defaultDatasetId")
|
||||
try:
|
||||
items = await plugins.get_dataset_items(dataset_id, limit=run.max_results)
|
||||
except (httpx.HTTPError, plugins.ApifyError, RuntimeError) as exc:
|
||||
raise HTTPException(status_code=502, detail=f"Apify dataset fetch failed: {exc}")
|
||||
normalized = [
|
||||
p
|
||||
for p in (plugins.normalize_profile(i) for i in items)
|
||||
if p and not plugins.is_excluded_profile(p)
|
||||
]
|
||||
job = await JobPosts.get_job_post_by_id(self.session, run.job_post_id)
|
||||
if job:
|
||||
job_fields = {
|
||||
"title": job.title,
|
||||
"requirements": job.requirements,
|
||||
"optional_skills": job.optional_skills,
|
||||
}
|
||||
for profile in normalized:
|
||||
profile["match_score"] = plugins.relevance_score(job_fields, profile)
|
||||
count = await TalentProfiles.upsert_from_items(
|
||||
self.session,
|
||||
job_post_id=run.job_post_id,
|
||||
run_id=run.id,
|
||||
normalized_items=normalized,
|
||||
)
|
||||
found_so_far = (run.profiles_found or 0) + count
|
||||
|
||||
# Thin results: broaden and keep the same run going instead of
|
||||
# settling for one lonely card. Each rung is a fresh actor run on
|
||||
# the same row; the frontend just sees "running" a while longer.
|
||||
if len(items) < min(10, run.max_results):
|
||||
broadened = plugins.broaden_actor_input(run.search_input or {})
|
||||
if broadened:
|
||||
try:
|
||||
started = await plugins.start_actor_run(broadened)
|
||||
except (httpx.HTTPError, plugins.ApifyError, RuntimeError):
|
||||
# Keep what we already found rather than failing the run.
|
||||
started = None
|
||||
if started:
|
||||
run = await TalentRuns.mark_rearmed(
|
||||
self.session,
|
||||
run.id,
|
||||
apify_run_id=started.get("id"),
|
||||
apify_dataset_id=started.get("defaultDatasetId"),
|
||||
search_input=broadened,
|
||||
found_so_far=found_so_far,
|
||||
)
|
||||
return serialize_talent_run(run)
|
||||
|
||||
run = await TalentRuns.mark_succeeded(
|
||||
self.session, run.id, profiles_found=found_so_far
|
||||
)
|
||||
return serialize_talent_run(run)
|
||||
|
||||
# failed / timed_out / aborted
|
||||
message = remote.get("statusMessage") or f"Apify run {remote.get('status')}"
|
||||
run = await TalentRuns.mark_failed(self.session, run.id, message, status=status)
|
||||
return serialize_talent_run(run)
|
||||
|
||||
async def fetch_runs(self, job_post_id):
|
||||
await self._get_job(job_post_id)
|
||||
rows, total = await TalentRuns.fetch_runs(self.session, job_post_id=job_post_id)
|
||||
return [serialize_talent_run(r) for r in rows], total
|
||||
|
||||
async def fetch_profiles(self, job_post_id, search=None, top=None, skip=0):
|
||||
await self._get_job(job_post_id)
|
||||
rows, total = await TalentProfiles.fetch_profiles(
|
||||
self.session, job_post_id=job_post_id, search=search, top=top, skip=skip
|
||||
)
|
||||
return [serialize_talent_profile(r) for r in rows], total
|
||||
|
||||
async def get_profile(self, profile_id):
|
||||
row = await TalentProfiles.get_profile_by_id(self.session, profile_id)
|
||||
if not row:
|
||||
raise HTTPException(status_code=404, detail="Talent profile not found")
|
||||
return serialize_talent_profile_detail(row)
|
||||
|
||||
async def delete_profile(self, profile_id):
|
||||
row = await TalentProfiles.soft_delete_profile(self.session, profile_id)
|
||||
if not row:
|
||||
raise HTTPException(status_code=404, detail="Talent profile not found")
|
||||
return {"id": str(row.id), "deleted": True}
|
||||
|
|
@ -35,6 +35,6 @@ def _hermetic_env(monkeypatch: pytest.MonkeyPatch) -> Iterator[None]:
|
|||
"""
|
||||
for name in list(os.environ):
|
||||
upper = name.upper()
|
||||
if upper.startswith(("OPENAI_", "ANTHROPIC_", "INBOX_TRIAGE_", "SCORING_", "MAX_")):
|
||||
if upper.startswith(("OPENAI_", "ANTHROPIC_", "INBOX_TRIAGE_", "SCORING_", "MAX_", "APIFY_")):
|
||||
monkeypatch.delenv(name, raising=False)
|
||||
yield
|
||||
|
|
|
|||
|
|
@ -0,0 +1,474 @@
|
|||
"""Unit tests for talent/plugins.py — the pure functions only.
|
||||
|
||||
No HTTP-call tests here, matching the Buffer adapter's precedent: the request
|
||||
helpers are thin httpx wrappers and the live smoke run covers them.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from talent import plugins
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- build_actor_input
|
||||
|
||||
def test_title_goes_to_the_facet_and_skills_to_the_query():
|
||||
result = plugins.build_actor_input(
|
||||
{
|
||||
"title": "Backend Engineer",
|
||||
"requirements": ["Python", "FastAPI", "PostgreSQL", "Docker", "AWS"],
|
||||
"location": "Berlin",
|
||||
},
|
||||
max_results=10,
|
||||
)
|
||||
assert result["currentJobTitles"] == ["Backend Engineer"]
|
||||
assert result["searchQuery"] == "Python FastAPI PostgreSQL"
|
||||
assert result["maxItems"] == 10
|
||||
assert result["locations"] == ["Berlin"]
|
||||
assert result["profileScraperMode"] == plugins.APIFY_PROFILE_MODE
|
||||
|
||||
|
||||
def test_actor_input_omits_locations_and_query_when_job_has_none():
|
||||
result = plugins.build_actor_input({"title": "Designer", "requirements": []}, max_results=5)
|
||||
assert "locations" not in result
|
||||
assert "searchQuery" not in result # the title facet alone carries the search
|
||||
assert result["currentJobTitles"] == ["Designer"]
|
||||
|
||||
|
||||
def test_non_geographic_locations_are_not_sent_as_filters():
|
||||
for value in ("Remote", "remote", "HYBRID", "Work From Home", " Onsite "):
|
||||
result = plugins.build_actor_input(
|
||||
{"title": "Dev", "requirements": [], "location": value}, max_results=5
|
||||
)
|
||||
assert "locations" not in result, value
|
||||
|
||||
|
||||
def test_remote_override_clears_the_location_filter_entirely():
|
||||
# An explicit "Remote" override means "no geography constraint" — it must
|
||||
# not be sent as a filter AND must not fall back to the job's location.
|
||||
result = plugins.build_actor_input(
|
||||
{"title": "Dev", "requirements": [], "location": "Berlin"},
|
||||
max_results=5,
|
||||
overrides={"location": "Remote"},
|
||||
)
|
||||
assert "locations" not in result
|
||||
|
||||
|
||||
def test_facet_title_is_capped_at_100_chars():
|
||||
result = plugins.build_actor_input(
|
||||
{"title": "X" * 300, "requirements": []}, max_results=5
|
||||
)
|
||||
assert result["currentJobTitles"] == ["X" * 100]
|
||||
|
||||
|
||||
def test_actor_input_overrides_win():
|
||||
result = plugins.build_actor_input(
|
||||
{"title": "Backend Engineer", "requirements": ["Python"], "location": "Berlin"},
|
||||
max_results=5,
|
||||
overrides={"keywords": "data engineer spark", "location": "Munich"},
|
||||
)
|
||||
assert result["searchQuery"] == "data engineer spark"
|
||||
assert result["locations"] == ["Munich"]
|
||||
assert result["currentJobTitles"] == ["Backend Engineer"]
|
||||
|
||||
|
||||
def test_actor_input_ignores_blank_requirement_entries():
|
||||
result = plugins.build_actor_input(
|
||||
{"title": "Dev", "requirements": [" ", "", "Go"]}, max_results=5
|
||||
)
|
||||
assert result["searchQuery"] == "Go"
|
||||
|
||||
|
||||
def test_sentence_requirements_stay_out_of_the_query():
|
||||
result = plugins.build_actor_input(
|
||||
{
|
||||
"title": "Amazon PPC",
|
||||
"requirements": [
|
||||
"2-5 years of experience managing Amazon PPC campaigns for e-commerce brands",
|
||||
"Strong hands-on experience with Amazon Ads, including Sponsored Products",
|
||||
],
|
||||
},
|
||||
max_results=5,
|
||||
)
|
||||
assert "searchQuery" not in result
|
||||
assert result["currentJobTitles"] == ["Amazon PPC"]
|
||||
|
||||
|
||||
def test_optional_skills_fill_in_when_requirements_are_prose():
|
||||
result = plugins.build_actor_input(
|
||||
{
|
||||
"title": "Amazon PPC",
|
||||
"requirements": ["Several sentences of prose describing years of experience required"],
|
||||
"optional_skills": ["Amazon Seller Central", "Helium 10", "PPC Bid Management", "Extra"],
|
||||
},
|
||||
max_results=5,
|
||||
)
|
||||
assert result["searchQuery"] == "Amazon Seller Central Helium 10 PPC Bid Management"
|
||||
|
||||
|
||||
def test_keyword_requirements_win_over_optional_skills():
|
||||
result = plugins.build_actor_input(
|
||||
{
|
||||
"title": "Dev",
|
||||
"requirements": ["Python", "FastAPI"],
|
||||
"optional_skills": ["Docker", "AWS"],
|
||||
},
|
||||
max_results=5,
|
||||
)
|
||||
assert result["searchQuery"] == "Python FastAPI Docker"
|
||||
|
||||
|
||||
def test_experience_range_selects_overlapping_buckets():
|
||||
assert plugins.years_of_experience_ids(3, 5) == ["3"]
|
||||
assert plugins.years_of_experience_ids(2, 4) == ["2", "3"]
|
||||
assert plugins.years_of_experience_ids(5, None) == ["3", "4", "5"]
|
||||
assert plugins.years_of_experience_ids(None, 1) == ["1", "2"]
|
||||
assert plugins.years_of_experience_ids(0, 60) == ["1", "2", "3", "4", "5"]
|
||||
assert plugins.years_of_experience_ids(None, None) == []
|
||||
|
||||
|
||||
def test_actor_input_carries_experience_filter():
|
||||
result = plugins.build_actor_input(
|
||||
{"title": "Dev", "requirements": [], "experience_min": 3, "experience_max": 5},
|
||||
max_results=5,
|
||||
)
|
||||
assert result["yearsOfExperienceIds"] == ["3"]
|
||||
no_exp = plugins.build_actor_input({"title": "Dev", "requirements": []}, max_results=5)
|
||||
assert "yearsOfExperienceIds" not in no_exp
|
||||
|
||||
|
||||
def test_actor_input_start_page():
|
||||
paged = plugins.build_actor_input({"title": "Dev"}, max_results=5, start_page=3)
|
||||
assert paged["startPage"] == 3
|
||||
first = plugins.build_actor_input({"title": "Dev"}, max_results=5, start_page=1)
|
||||
assert "startPage" not in first # page 1 is the actor default; keep input stable
|
||||
capped = plugins.build_actor_input({"title": "Dev"}, max_results=5, start_page=999)
|
||||
assert capped["startPage"] == 100
|
||||
|
||||
|
||||
def test_actor_input_excludes_own_company_urls():
|
||||
result = plugins.build_actor_input({"title": "Dev"}, max_results=5)
|
||||
assert result["excludeCurrentCompanies"] == plugins.APIFY_EXCLUDE_COMPANY_URLS
|
||||
assert any("utopiadeals" in u for u in result["excludeCurrentCompanies"])
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- own-company filter
|
||||
|
||||
def test_current_utopia_employees_are_excluded():
|
||||
assert plugins.is_excluded_profile({"current_company": "Utopia Brands"})
|
||||
assert plugins.is_excluded_profile({"current_company": "utopia deals"})
|
||||
assert plugins.is_excluded_profile({"current_company": "Utopia Brands Pakistan (Pvt) Ltd"})
|
||||
|
||||
|
||||
def test_other_companies_and_former_employees_pass():
|
||||
assert not plugins.is_excluded_profile({"current_company": "Acme"})
|
||||
assert not plugins.is_excluded_profile({"current_company": None})
|
||||
assert not plugins.is_excluded_profile({})
|
||||
# Headline mentioning Utopia does NOT exclude someone whose current company
|
||||
# is elsewhere (e.g. "ex-Utopia Deals, now at Acme").
|
||||
assert not plugins.is_excluded_profile(
|
||||
{"current_company": "Acme", "headline": "ex-Utopia Deals engineer"}
|
||||
)
|
||||
|
||||
|
||||
def test_headline_is_the_fallback_when_company_is_missing():
|
||||
assert plugins.is_excluded_profile(
|
||||
{"current_company": None, "headline": "SEO Executive at Utopia Deals"}
|
||||
)
|
||||
assert not plugins.is_excluded_profile(
|
||||
{"current_company": None, "headline": "Backend Engineer"}
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- local_status
|
||||
|
||||
def test_every_known_apify_status_maps():
|
||||
assert plugins.local_status("READY") == "running"
|
||||
assert plugins.local_status("RUNNING") == "running"
|
||||
assert plugins.local_status("TIMING-OUT") == "running"
|
||||
assert plugins.local_status("ABORTING") == "running"
|
||||
assert plugins.local_status("SUCCEEDED") == "succeeded"
|
||||
assert plugins.local_status("FAILED") == "failed"
|
||||
assert plugins.local_status("TIMED-OUT") == "timed_out"
|
||||
assert plugins.local_status("ABORTED") == "aborted"
|
||||
|
||||
|
||||
def test_unknown_and_missing_statuses_stay_running():
|
||||
assert plugins.local_status("SOMETHING-NEW") == "running"
|
||||
assert plugins.local_status(None) == "running"
|
||||
assert plugins.local_status("") == "running"
|
||||
|
||||
|
||||
def test_terminal_statuses_are_the_terminal_local_values():
|
||||
assert plugins.TERMINAL_STATUSES == {"succeeded", "failed", "timed_out", "aborted"}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- normalize_linkedin_url
|
||||
|
||||
def test_url_normalization_canonicalizes():
|
||||
expected = "https://www.linkedin.com/in/jane-doe"
|
||||
assert plugins.normalize_linkedin_url("https://www.linkedin.com/in/Jane-Doe/") == expected
|
||||
assert plugins.normalize_linkedin_url("http://www.LinkedIn.com/in/jane-doe?src=x#top") == expected
|
||||
assert plugins.normalize_linkedin_url("www.linkedin.com/in/jane-doe") == expected
|
||||
|
||||
|
||||
def test_url_normalization_rejects_non_linkedin():
|
||||
assert plugins.normalize_linkedin_url("https://twitter.com/janedoe") is None
|
||||
assert plugins.normalize_linkedin_url("") is None
|
||||
assert plugins.normalize_linkedin_url(None) is None
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- normalize_profile
|
||||
|
||||
# Shape observed live from harvestapi~linkedin-profile-search (Full mode).
|
||||
RICH_ITEM = {
|
||||
"linkedinUrl": "https://www.linkedin.com/in/sachinsharma31261",
|
||||
"publicIdentifier": "sachinsharma31261",
|
||||
"firstName": "Sachin",
|
||||
"lastName": "Sharma",
|
||||
"headline": "Software Engineer @ Lucid Motors",
|
||||
"about": "Staff Software Engineer with 10 years of experience.",
|
||||
"location": {"linkedinText": "San Jose, California, United States"},
|
||||
"photo": "https://media.licdn.com/photo.jpg",
|
||||
"currentPosition": {"title": "Lead Software Engineer", "companyName": "Lucid Motors"},
|
||||
"experience": [{"title": "Lead Software Engineer", "companyName": "Lucid Motors"}],
|
||||
"skills": [{"name": "Java"}, {"name": "Python"}, {"name": "Java"}],
|
||||
}
|
||||
|
||||
|
||||
def test_rich_item_normalizes_every_card_field():
|
||||
profile = plugins.normalize_profile(RICH_ITEM)
|
||||
assert profile["linkedin_url"] == "https://www.linkedin.com/in/sachinsharma31261"
|
||||
assert profile["public_id"] == "sachinsharma31261"
|
||||
assert profile["full_name"] == "Sachin Sharma"
|
||||
assert profile["headline"] == "Software Engineer @ Lucid Motors"
|
||||
assert profile["location"] == "San Jose, California, United States"
|
||||
assert profile["current_title"] == "Lead Software Engineer"
|
||||
assert profile["current_company"] == "Lucid Motors"
|
||||
assert profile["avatar_url"] == "https://media.licdn.com/photo.jpg"
|
||||
assert profile["summary"] == "Staff Software Engineer with 10 years of experience."
|
||||
assert profile["skills"] == ["Java", "Python"] # dict entries, deduped
|
||||
assert profile["raw"] is RICH_ITEM
|
||||
|
||||
|
||||
def test_skills_accept_plain_strings_and_prefer_top_skills():
|
||||
profile = plugins.normalize_profile({
|
||||
"linkedinUrl": "https://linkedin.com/in/x",
|
||||
"topSkills": ["Go", "Rust"],
|
||||
"skills": [{"name": "Ignored"}],
|
||||
})
|
||||
assert profile["skills"] == ["Go", "Rust"]
|
||||
none = plugins.normalize_profile({"linkedinUrl": "https://linkedin.com/in/y"})
|
||||
assert none["skills"] == []
|
||||
|
||||
|
||||
def test_minimal_item_still_normalizes():
|
||||
profile = plugins.normalize_profile(
|
||||
{"url": "https://linkedin.com/in/someone", "name": "Some One"}
|
||||
)
|
||||
assert profile["linkedin_url"] == "https://linkedin.com/in/someone"
|
||||
assert profile["full_name"] == "Some One"
|
||||
assert profile["headline"] is None
|
||||
assert profile["avatar_url"] is None
|
||||
|
||||
|
||||
def test_item_without_linkedin_url_is_skipped_not_fatal():
|
||||
assert plugins.normalize_profile({"name": "No Url"}) is None
|
||||
assert plugins.normalize_profile({"url": "https://example.com/x"}) is None
|
||||
assert plugins.normalize_profile("not a dict") is None
|
||||
|
||||
|
||||
def test_location_accepts_plain_string():
|
||||
profile = plugins.normalize_profile(
|
||||
{"linkedinUrl": "https://linkedin.com/in/x", "location": "Greater St. Louis"}
|
||||
)
|
||||
assert profile["location"] == "Greater St. Louis"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- broadening ladder
|
||||
|
||||
def test_broadening_ladder_relaxes_one_constraint_per_rung():
|
||||
original = {
|
||||
"currentJobTitles": ["Amazon PPC"],
|
||||
"searchQuery": "Amazon DSP experience",
|
||||
"locations": ["Karachi, Pakistan"],
|
||||
"yearsOfExperienceIds": ["2", "3"],
|
||||
"maxItems": 25,
|
||||
"profileScraperMode": "Full",
|
||||
}
|
||||
rung1 = plugins.broaden_actor_input(original)
|
||||
assert "searchQuery" not in rung1
|
||||
assert rung1["currentJobTitles"] == ["Amazon PPC"]
|
||||
assert rung1["locations"] == ["Karachi, Pakistan"] # never relaxed
|
||||
assert rung1["yearsOfExperienceIds"] == ["2", "3"] # never relaxed
|
||||
|
||||
rung2 = plugins.broaden_actor_input(rung1)
|
||||
assert "currentJobTitles" not in rung2
|
||||
assert rung2["searchQuery"] == "Amazon PPC" # title as keywords
|
||||
assert rung2["locations"] == ["Karachi, Pakistan"]
|
||||
|
||||
assert plugins.broaden_actor_input(rung2) is None # exhausted
|
||||
|
||||
|
||||
def test_broadening_does_not_mutate_the_original_input():
|
||||
original = {"currentJobTitles": ["Dev"], "searchQuery": "Python"}
|
||||
plugins.broaden_actor_input(original)
|
||||
assert original == {"currentJobTitles": ["Dev"], "searchQuery": "Python"}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- relevance score
|
||||
|
||||
AI_JOB = {
|
||||
"title": "AI Engineer",
|
||||
"requirements": ["Python", "FastAPI", "PostgreSQL"],
|
||||
"optional_skills": [],
|
||||
}
|
||||
|
||||
|
||||
def test_actual_ai_engineer_outranks_keyword_stuffed_full_stack():
|
||||
# The live case that motivated the phrase rule: Khawar's keyword-stuffed
|
||||
# headline carries every hot token ("AI/ML Engineer | Python | FastAPI |
|
||||
# ...") but his title is Full Stack; Shaheer's title IS "AI Engineer".
|
||||
full_stack = {
|
||||
"current_title": "Sr. Full Stack Engineer",
|
||||
"headline": (
|
||||
"Senior Software Engineer| Senior Full Stack Engineer | AI/ML Engineer "
|
||||
"| Python | FastAPI | Django | React| LLMs | RAG | Agentic AI | AWS"
|
||||
),
|
||||
"skills": ["Python (Programming Language)", "JavaScript", "React.js"],
|
||||
"summary": "Senior Software Engineer delivering web applications with PostgreSQL.",
|
||||
}
|
||||
ai_engineer = {
|
||||
"current_title": "AI Engineer",
|
||||
"headline": "AI Engineer @ EmpireOne | Building Production LLM Systems",
|
||||
"skills": ["Keras", "Docker", "FastAPI", "PostgreSQL", "Python"],
|
||||
"summary": "Machine Learning and Data Science.",
|
||||
}
|
||||
weak = plugins.relevance_score(AI_JOB, full_stack)
|
||||
strong = plugins.relevance_score(AI_JOB, ai_engineer)
|
||||
assert strong > weak
|
||||
assert strong >= 55 # exact title phrase at minimum
|
||||
assert weak <= 80 # scattered tokens cap at 35 + full skills 45
|
||||
|
||||
|
||||
def test_relevance_score_bounds_and_empty_profile():
|
||||
perfect = plugins.relevance_score(AI_JOB, {
|
||||
"current_title": "AI Engineer",
|
||||
"skills": ["Python", "FastAPI", "PostgreSQL"],
|
||||
})
|
||||
assert perfect == 100
|
||||
assert plugins.relevance_score(AI_JOB, {}) == 0
|
||||
assert plugins.relevance_score({"title": "", "requirements": []}, {"headline": "x"}) == 0
|
||||
|
||||
|
||||
def test_skills_overlap_is_graded_not_all_or_nothing():
|
||||
# A single unmatched niche term must not zero the whole skills component:
|
||||
# that put an entire live pool on exactly 60.
|
||||
job = {"title": "PPC", "requirements": ["Amazon Seller Central"], "optional_skills": []}
|
||||
full = plugins.relevance_score(job, {"skills": ["Amazon Seller Central"], "current_title": "PPC"})
|
||||
partial = plugins.relevance_score(job, {"skills": ["Amazon"], "current_title": "PPC"})
|
||||
none = plugins.relevance_score(job, {"skills": ["Photoshop"], "current_title": "PPC"})
|
||||
assert full == 100
|
||||
assert none == 55 # title only
|
||||
assert none < partial < full # 1 of 3 tokens matched sits in between
|
||||
|
||||
|
||||
def test_prose_requirements_still_differentiate_profiles():
|
||||
# The Amazon PPC case: prose requirements yielded one niche term and every
|
||||
# sourced profile scored identically. Graded token overlap must spread them.
|
||||
job = {
|
||||
"title": "Amazon PPC",
|
||||
"requirements": [
|
||||
"2-5 years of experience managing Amazon PPC campaigns for e-commerce brands",
|
||||
"Strong hands-on experience with Amazon Ads, including Sponsored Products",
|
||||
],
|
||||
"optional_skills": [],
|
||||
}
|
||||
rich = plugins.relevance_score(job, {
|
||||
"current_title": "Amazon PPC Manager",
|
||||
"skills": ["Amazon PPC", "PPC Bid Management", "Amazon Listing Optimization"],
|
||||
"summary": "Managing Amazon Ads campaigns, Sponsored Products and Sponsored Display for e-commerce brands.",
|
||||
})
|
||||
thin = plugins.relevance_score(job, {
|
||||
"current_title": "Amazon PPC Specialist",
|
||||
"skills": [],
|
||||
"summary": "",
|
||||
})
|
||||
assert rich > thin >= 55
|
||||
assert rich - thin >= 15 # a real spread, not a flat pool
|
||||
|
||||
|
||||
def test_headline_phrase_scores_below_title_phrase():
|
||||
job = {"title": "AI Engineer", "requirements": [], "optional_skills": []}
|
||||
in_title = plugins.relevance_score(job, {"current_title": "AI Engineer"})
|
||||
in_headline = plugins.relevance_score(job, {"current_title": "Developer", "headline": "AI Engineer at Acme"})
|
||||
scattered = plugins.relevance_score(job, {"current_title": "Engineer", "headline": "Agentic AI | Python"})
|
||||
assert in_title == 55
|
||||
assert in_headline == 45
|
||||
assert scattered == 35 # both tokens present but never as the phrase
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- detail extraction
|
||||
|
||||
RAW_WITH_HISTORY = {
|
||||
"experience": [
|
||||
{
|
||||
"position": "Freelance",
|
||||
"companyName": "Upwork",
|
||||
"employmentType": "Self-employed",
|
||||
"location": "Rawalpindi, Punjab, Pakistan",
|
||||
"duration": "1 yr 3 mos",
|
||||
"description": None,
|
||||
"skills": ["Amazon Seller Central", "Amazon PPC"],
|
||||
"startDate": {"month": "Jun", "year": 2025, "text": "Jun 2025"},
|
||||
"endDate": {"text": "Present"},
|
||||
},
|
||||
"not a dict",
|
||||
],
|
||||
"education": [
|
||||
{
|
||||
"schoolName": "Modern Public School - Pakistan",
|
||||
"degree": "Intermediate",
|
||||
"fieldOfStudy": "Computer Science",
|
||||
"period": "May 2020 - Jun 2022",
|
||||
},
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
def test_experience_extraction_matches_live_shape():
|
||||
entries = plugins.extract_experience(RAW_WITH_HISTORY)
|
||||
assert len(entries) == 1
|
||||
entry = entries[0]
|
||||
assert entry["title"] == "Freelance"
|
||||
assert entry["company"] == "Upwork"
|
||||
assert entry["employment_type"] == "Self-employed"
|
||||
assert entry["duration"] == "1 yr 3 mos"
|
||||
assert entry["period"] == "Jun 2025 – Present"
|
||||
assert entry["description"] is None
|
||||
assert entry["skills"] == ["Amazon Seller Central", "Amazon PPC"]
|
||||
|
||||
|
||||
def test_education_extraction_matches_live_shape():
|
||||
entries = plugins.extract_education(RAW_WITH_HISTORY)
|
||||
assert entries == [{
|
||||
"school": "Modern Public School - Pakistan",
|
||||
"degree": "Intermediate",
|
||||
"field": "Computer Science",
|
||||
"period": "May 2020 - Jun 2022",
|
||||
}]
|
||||
|
||||
|
||||
def test_history_extraction_tolerates_empty_raw():
|
||||
assert plugins.extract_experience({}) == []
|
||||
assert plugins.extract_education({}) == []
|
||||
assert plugins.extract_experience(None) == []
|
||||
assert plugins.extract_education(None) == []
|
||||
|
||||
|
||||
def test_company_from_nested_experience_company_dict():
|
||||
profile = plugins.normalize_profile({
|
||||
"linkedinUrl": "https://linkedin.com/in/x",
|
||||
"experience": [{"title": "Engineer", "company": {"name": "Acme"}}],
|
||||
})
|
||||
assert profile["current_title"] == "Engineer"
|
||||
assert profile["current_company"] == "Acme"
|
||||
|
|
@ -39,6 +39,7 @@ class PermissionModule(str, Enum):
|
|||
SETTINGS = "settings"
|
||||
RBAC_USERS = "rbac_users"
|
||||
TASKS = "tasks"
|
||||
TALENT = "talent"
|
||||
|
||||
|
||||
class PermissionAction(str, Enum):
|
||||
|
|
@ -165,6 +166,14 @@ class PermissionTag(str, Enum):
|
|||
TASKS_EXPORT = "tasks.export"
|
||||
TASKS_MANAGE = "tasks.manage"
|
||||
TASKS_CONFIGURE = "tasks.configure"
|
||||
TALENT_VIEW = "talent.view"
|
||||
TALENT_CREATE = "talent.create"
|
||||
TALENT_EDIT = "talent.edit"
|
||||
TALENT_DELETE = "talent.delete"
|
||||
TALENT_APPROVE = "talent.approve"
|
||||
TALENT_EXPORT = "talent.export"
|
||||
TALENT_MANAGE = "talent.manage"
|
||||
TALENT_CONFIGURE = "talent.configure"
|
||||
|
||||
|
||||
def _assert_vocabulary_complete() -> None:
|
||||
|
|
|
|||
|
|
@ -46,6 +46,10 @@ x-backend-env: &backend-env
|
|||
REDIS_URL: redis://redis:6379/0
|
||||
EMAIL_URL: http://host.docker.internal:5000
|
||||
BACKEND_URL: http://backend-api:8000
|
||||
# Apify talent sourcing: interpolated from the shell or a root .env, so a token
|
||||
# kept at repo root (APIFY_TOKEN) reaches containers without duplicating it
|
||||
# into backend/.env.
|
||||
APIFY_API_TOKEN: ${APIFY_API_TOKEN:-${APIFY_TOKEN:-}}
|
||||
|
||||
# The one shared folder. Every process that decodes, scores or serves a CV reads and
|
||||
# writes the same host directory, so a file written by the API is the same file the
|
||||
|
|
|
|||
|
|
@ -23,8 +23,8 @@
|
|||
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
|
||||
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=Belleza&family=Inter:wght@400;500;600;700;800&display=swap" />
|
||||
<link rel="icon" href="data:image/svg+xml,<svg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 100 100'><rect width='100' height='100' rx='22' fill='%23004d43'/><g transform='translate(14 32) scale(0.72)'><path d='M100 3.65C97.86 20.99 91.89 43.03 79.48 55.33 76 58.77 71.84 61.46 66.96 62.14 50.4 64.46 41.84 47.5 29.07 42.7 21.85 39.98 14.5 42.02 9.66 47.95 6.54 51.78 4.49 56.35 2.97 61.13 2.41 61.64 0.97 61.66 0 61.31L0 0.13C1.05 0 2.27 0.02 3.09 0.28 14.9 15.86 26.77 30.82 40.15 45.28L60.79 24.7C67.38 18.22 74.41 12.74 82.59 8.51 88.11 5.83 93.64 3.93 100 3.65Z' fill='%23ceff71'/></g></svg>" />
|
||||
<script type="module" crossorigin src="/assets/index-CpVGHhXU.js"></script>
|
||||
<link rel="stylesheet" crossorigin href="/assets/index-CP8rR_Xd.css">
|
||||
<script type="module" crossorigin src="/assets/index-DoHV5rbL.js"></script>
|
||||
<link rel="stylesheet" crossorigin href="/assets/index-BoTJFWhm.css">
|
||||
</head>
|
||||
<body>
|
||||
<div id="root"></div>
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
/**
|
||||
* Render smoke test — mounts all 27 routes (23 app + 4 auth) into jsdom and
|
||||
* fails on any thrown error, console.error, or empty render.
|
||||
* Render smoke test — mounts every route (all app routes + 4 auth) into jsdom
|
||||
* and fails on any thrown error, console.error, or empty render.
|
||||
*
|
||||
* npm run smoke
|
||||
*
|
||||
|
|
@ -48,9 +48,9 @@ dom.window.matchMedia = () => ({
|
|||
dom.window.HTMLCanvasElement.prototype.getContext = () =>
|
||||
new Proxy({}, { get: (_, k) => (k === 'canvas' ? {} : () => {}) })
|
||||
|
||||
// A signed-in session holding all 104 permissions, so no route is gated away.
|
||||
// A signed-in session holding every permission tag, so no route is gated away.
|
||||
const MODULES = ['dashboard', 'inbox', 'jobs', 'candidates', 'pipeline', 'interviews', 'assessments',
|
||||
'offers', 'reports', 'analytics', 'job_board', 'settings', 'rbac_users']
|
||||
'offers', 'reports', 'analytics', 'job_board', 'settings', 'rbac_users', 'tasks', 'talent']
|
||||
const ACTIONS = ['view', 'create', 'edit', 'delete', 'approve', 'export', 'manage', 'configure']
|
||||
const permissions = MODULES.flatMap((m) => ACTIONS.map((a) => `${m}.${a}`))
|
||||
|
||||
|
|
@ -96,9 +96,11 @@ console.error = (...args) => {
|
|||
}
|
||||
|
||||
let failed = 0
|
||||
let total = 0
|
||||
try {
|
||||
const mod = await import(pathToFileURL(outFile).href)
|
||||
mod.boot()
|
||||
total = mod.ALL_ROUTES.length
|
||||
|
||||
for (const path of mod.ALL_ROUTES) {
|
||||
errors.length = 0
|
||||
|
|
@ -127,5 +129,5 @@ try {
|
|||
rmSync(outDir, { recursive: true, force: true })
|
||||
}
|
||||
|
||||
console.log(failed ? `\n${failed}/27 routes FAILED` : `\nAll 27 routes rendered clean`)
|
||||
console.log(failed ? `\n${failed}/${total} routes FAILED` : `\nAll ${total} routes rendered clean`)
|
||||
process.exit(failed ? 1 : 0)
|
||||
|
|
|
|||
|
|
@ -25,6 +25,7 @@ const SCREENS = {
|
|||
import: lazy(() => import('./screens/CvImport')),
|
||||
jobboard: lazy(() => import('./screens/JobBoard')),
|
||||
recruiterhub: lazy(() => import('./screens/RecruiterHub')),
|
||||
talent: lazy(() => import('./screens/Talent')),
|
||||
tasks: lazy(() => import('./screens/Tasks')),
|
||||
aiassistant: lazy(() => import('./screens/AiAssistant')),
|
||||
interviews: lazy(() => import('./screens/Interviews')),
|
||||
|
|
|
|||
|
|
@ -33,6 +33,7 @@ import Pipeline from '../screens/Pipeline'
|
|||
import CvImport from '../screens/CvImport'
|
||||
import JobBoard from '../screens/JobBoard'
|
||||
import RecruiterHub from '../screens/RecruiterHub'
|
||||
import Talent from '../screens/Talent'
|
||||
import Tasks from '../screens/Tasks'
|
||||
import AiAssistant from '../screens/AiAssistant'
|
||||
import Interviews from '../screens/Interviews'
|
||||
|
|
@ -51,7 +52,7 @@ import Help from '../screens/Help'
|
|||
const SCREENS = {
|
||||
dashboard: Dashboard, inbox: Inbox, matching: Matching, jobs: Jobs, candidates: Candidates,
|
||||
talentpool: TalentPool, pipeline: Pipeline, import: CvImport, jobboard: JobBoard,
|
||||
recruiterhub: RecruiterHub, tasks: Tasks, aiassistant: AiAssistant,
|
||||
recruiterhub: RecruiterHub, talent: Talent, tasks: Tasks, aiassistant: AiAssistant,
|
||||
interviews: Interviews, assessments: Assessments, offers: Offers,
|
||||
managers: Managers, calendar: Calendar, reports: Reports, analytics: Analytics,
|
||||
aistudio: AiStudio, notifications: Notifications, rbac: Rbac,
|
||||
|
|
|
|||
|
|
@ -0,0 +1,117 @@
|
|||
/* ============================================================
|
||||
talent.js — LinkedIn talent sourcing endpoints (backend/talent/app.py).
|
||||
|
||||
A "run" is one paid Apify actor search for a job post; profiles are the
|
||||
deduped people those runs found. Same conventions as candidates.js: one
|
||||
named export per endpoint, no hooks, camelCase params mapped to snake_case
|
||||
at the call boundary, and every function returns the parsed
|
||||
{data, total, status_code} envelope.
|
||||
============================================================ */
|
||||
|
||||
import { request } from '../lib/apiClient'
|
||||
|
||||
const TERMINAL = new Set(['succeeded', 'failed', 'timed_out', 'aborted'])
|
||||
|
||||
/** Start a paid sourcing run for a job. Needs talent.create. 409s while one is running. */
|
||||
export function startRun(jobPostId, { maxResults, location, keywords } = {}) {
|
||||
return request('/talent/runs/start', {
|
||||
method: 'POST',
|
||||
params: { job_post_id: jobPostId },
|
||||
body: { max_results: maxResults, location, keywords },
|
||||
})
|
||||
}
|
||||
|
||||
/**
|
||||
* Poll target. Needs talent.view. When Apify reports the run finished, THIS
|
||||
* call persists the found profiles server-side before answering — so polling
|
||||
* it is what completes a run, even after a page reload.
|
||||
*/
|
||||
export function getRunStatus(runId) {
|
||||
return request('/talent/runs/status', { params: { run_id: runId } })
|
||||
}
|
||||
|
||||
/** Run history for a job, newest first. Needs talent.view. */
|
||||
export function listRuns(jobPostId) {
|
||||
return request('/talent/runs/fetch', { params: { job_post_id: jobPostId } })
|
||||
}
|
||||
|
||||
/** Sourced profiles for a job, most recently seen first. Needs talent.view. */
|
||||
export function listProfiles({ jobId, search, top, skip } = {}) {
|
||||
return request('/talent/profiles/fetch', {
|
||||
params: { job_post_id: jobId, search, top, skip },
|
||||
})
|
||||
}
|
||||
|
||||
/** One profile with employment/education history unpacked. Needs talent.view. */
|
||||
export function getProfile(profileId) {
|
||||
return request('/talent/profiles/fetch_by_id', { params: { profile_id: profileId } })
|
||||
}
|
||||
|
||||
/** Dismiss a profile (soft delete; re-runs will not resurrect it). Needs talent.delete. */
|
||||
export function deleteProfile(profileId) {
|
||||
return request('/talent/profiles/delete', {
|
||||
method: 'DELETE',
|
||||
params: { profile_id: profileId },
|
||||
})
|
||||
}
|
||||
|
||||
export function isTerminalRun(status) {
|
||||
return TERMINAL.has(status)
|
||||
}
|
||||
|
||||
export function toRunView(row) {
|
||||
return {
|
||||
id: row.id,
|
||||
jobId: row.job_post_id,
|
||||
status: row.status,
|
||||
maxResults: row.max_results ?? null,
|
||||
profilesFound: row.profiles_found ?? 0,
|
||||
error: row.apify_error ?? null,
|
||||
startedAt: row.started_at ? new Date(row.started_at) : null,
|
||||
finishedAt: row.finished_at ? new Date(row.finished_at) : null,
|
||||
createdAt: row.created_at ? new Date(row.created_at) : null,
|
||||
isTerminal: TERMINAL.has(row.status),
|
||||
}
|
||||
}
|
||||
|
||||
export function toProfileView(row) {
|
||||
return {
|
||||
id: row.id,
|
||||
jobId: row.job_post_id,
|
||||
name: row.full_name,
|
||||
headline: row.headline ?? null,
|
||||
location: row.location ?? null,
|
||||
currentTitle: row.current_title ?? null,
|
||||
currentCompany: row.current_company ?? null,
|
||||
avatarUrl: row.avatar_url ?? null,
|
||||
linkedinUrl: row.linkedin_url,
|
||||
publicId: row.public_id ?? null,
|
||||
summary: row.summary ?? null,
|
||||
skills: Array.isArray(row.skills) ? row.skills : [],
|
||||
matchScore: row.match_score ?? null,
|
||||
lastSeenAt: row.last_seen_at ? new Date(row.last_seen_at) : null,
|
||||
}
|
||||
}
|
||||
|
||||
export function toProfileDetailView(row) {
|
||||
return {
|
||||
...toProfileView(row),
|
||||
firstSeenAt: row.first_seen_at ? new Date(row.first_seen_at) : null,
|
||||
experience: (row.experience ?? []).map((e) => ({
|
||||
title: e.title,
|
||||
company: e.company,
|
||||
employmentType: e.employment_type,
|
||||
location: e.location,
|
||||
duration: e.duration,
|
||||
period: e.period,
|
||||
description: e.description,
|
||||
skills: Array.isArray(e.skills) ? e.skills : [],
|
||||
})),
|
||||
education: (row.education ?? []).map((e) => ({
|
||||
school: e.school,
|
||||
degree: e.degree,
|
||||
field: e.field,
|
||||
period: e.period,
|
||||
})),
|
||||
}
|
||||
}
|
||||
|
|
@ -28,6 +28,7 @@ export const ROUTES = [
|
|||
{ path: 'import', title: 'CV Import', icon: 'upload', group: 'Recruiting', permission: 'candidates.create' },
|
||||
{ path: 'jobboard', title: 'Job Board', icon: 'layers', group: 'Recruiting', permission: 'job_board.view' },
|
||||
{ path: 'recruiterhub', title: 'Recruiter Hub', icon: 'check-circle', group: 'Recruiting', permission: 'analytics.view' },
|
||||
{ path: 'talent', title: 'Talent', icon: 'user-plus', group: 'Recruiting', permission: 'talent.view' },
|
||||
{ path: 'tasks', title: 'Tasks', icon: 'check-square', group: 'Recruiting', permission: 'tasks.view', badge: 'tasks' },
|
||||
{ path: 'aiassistant', title: 'AI Assistant', icon: 'sparkles', group: 'Recruiting', permission: null, tag: 'AI' },
|
||||
|
||||
|
|
|
|||
|
|
@ -15,14 +15,15 @@
|
|||
|
||||
export const MODULES = [
|
||||
'dashboard', 'inbox', 'jobs', 'candidates', 'pipeline', 'interviews', 'assessments',
|
||||
'offers', 'reports', 'analytics', 'job_board', 'settings', 'rbac_users',
|
||||
'offers', 'reports', 'analytics', 'job_board', 'settings', 'rbac_users', 'tasks',
|
||||
'talent',
|
||||
]
|
||||
|
||||
export const ACTIONS = [
|
||||
'view', 'create', 'edit', 'delete', 'approve', 'export', 'manage', 'configure',
|
||||
]
|
||||
|
||||
/** All 104 `module.action` tags. */
|
||||
/** All `module.action` tags (modules x actions cross-product). */
|
||||
export const ALL_TAGS = MODULES.flatMap((m) => ACTIONS.map((a) => `${m}.${a}`))
|
||||
|
||||
/**
|
||||
|
|
|
|||
|
|
@ -64,6 +64,13 @@ export const qk = {
|
|||
all: () => ['jobs'],
|
||||
list: (p = {}) => ['jobs', 'list', p],
|
||||
},
|
||||
talent: {
|
||||
all: () => ['talent'],
|
||||
runs: (jobId) => ['talent', 'runs', jobId],
|
||||
run: (runId) => ['talent', 'run', runId],
|
||||
profiles: (p = {}) => ['talent', 'profiles', p],
|
||||
profile: (id) => ['talent', 'profile', id],
|
||||
},
|
||||
candidates: {
|
||||
all: () => ['candidates'],
|
||||
list: (p = {}) => ['candidates', 'list', p],
|
||||
|
|
|
|||
|
|
@ -0,0 +1,598 @@
|
|||
/* ============================================================
|
||||
Talent — LinkedIn talent sourcing per job (backend/talent/, Apify).
|
||||
|
||||
Pick a job, start a paid actor search, watch the run, browse the profiles.
|
||||
The status poll is what persists results server-side: the backend fetches
|
||||
the Apify dataset the first time it sees the run SUCCEEDED, so reloading
|
||||
mid-run loses nothing — the screen re-adopts the newest unfinished run and
|
||||
keeps polling. Profiles are deduped per job across re-runs by LinkedIn URL.
|
||||
============================================================ */
|
||||
|
||||
import { useEffect, useMemo, useRef, useState } from 'react'
|
||||
import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query'
|
||||
|
||||
import Modal from '../ui/Modal'
|
||||
import { Badge, EmptyState, Icon } from '../ui/primitives'
|
||||
import { useToast } from '../ui/Toast'
|
||||
import { qk } from '../lib/queryKeys'
|
||||
import { friendlyAuthError } from '../lib/errors'
|
||||
import * as candidatesApi from '../api/candidates'
|
||||
import * as talentApi from '../api/talent'
|
||||
import { avatarColor, fmtDate, initials as initialsOf } from '../data/seed'
|
||||
|
||||
const RUN_BADGE = {
|
||||
pending: ['b-blue', 'Starting…'],
|
||||
running: ['b-blue', 'Sourcing…'],
|
||||
succeeded: ['b-green', 'Completed'],
|
||||
failed: ['b-red', 'Failed'],
|
||||
timed_out: ['b-red', 'Timed out'],
|
||||
aborted: ['b-amber', 'Aborted'],
|
||||
}
|
||||
|
||||
async function fetchJobs() {
|
||||
const res = await candidatesApi.listJobs()
|
||||
const rows = Array.isArray(res?.data) ? res.data : []
|
||||
return rows.map((row) => ({ id: row.id, title: row.title, location: row.location }))
|
||||
}
|
||||
|
||||
/* Where to source from. Pakistan is the company's hub (Karachi and Lahore
|
||||
offices today), so those lead the list; "Anywhere" clears the geography
|
||||
filter server-side (useful for remote roles); CUSTOM reveals a free-text
|
||||
input for anything else. */
|
||||
const CUSTOM_LOCATION = '__custom__'
|
||||
const LOCATION_OPTIONS = [
|
||||
{ value: 'Karachi, Pakistan', label: 'Karachi' },
|
||||
{ value: 'Lahore, Pakistan', label: 'Lahore' },
|
||||
{ value: 'Pakistan', label: 'Pakistan — country-wide' },
|
||||
{ value: 'Anywhere', label: 'Anywhere (no location filter)' },
|
||||
{ value: CUSTOM_LOCATION, label: 'Custom location…' },
|
||||
]
|
||||
|
||||
/* Work arrangements are not geographies — mirror of the backend list. */
|
||||
const NON_GEOGRAPHIC = new Set([
|
||||
'remote', 'hybrid', 'onsite', 'on-site', 'on site',
|
||||
'anywhere', 'flexible', 'wfh', 'work from home',
|
||||
])
|
||||
|
||||
/** Dropdown default for a job: its own city when it has one, else the hub. */
|
||||
function defaultLocationFor(job) {
|
||||
const loc = (job?.location || '').trim()
|
||||
if (!loc || NON_GEOGRAPHIC.has(loc.toLowerCase())) {
|
||||
// Remote/unspecified posts still source from the hub by default; the
|
||||
// recruiter can widen to country-wide or Anywhere from the dropdown.
|
||||
return { choice: 'Pakistan', custom: '' }
|
||||
}
|
||||
const match = LOCATION_OPTIONS.find(
|
||||
(o) => o.value !== CUSTOM_LOCATION && o.value.toLowerCase().startsWith(loc.toLowerCase()),
|
||||
)
|
||||
if (match) return { choice: match.value, custom: '' }
|
||||
return { choice: CUSTOM_LOCATION, custom: loc }
|
||||
}
|
||||
|
||||
function ProfileAvatar({ name, url }) {
|
||||
const [broken, setBroken] = useState(false)
|
||||
if (url && !broken) {
|
||||
return (
|
||||
<img
|
||||
className="avatar avatar-lg"
|
||||
src={url}
|
||||
alt={name || 'Profile photo'}
|
||||
referrerPolicy="no-referrer"
|
||||
style={{ objectFit: 'cover', padding: 0 }}
|
||||
onError={() => setBroken(true)}
|
||||
/>
|
||||
)
|
||||
}
|
||||
return (
|
||||
<span className="avatar avatar-lg" style={{ background: avatarColor(name || '?') }}>
|
||||
{initialsOf(name || '?')}
|
||||
</span>
|
||||
)
|
||||
}
|
||||
|
||||
/** Big centered loader: the ai-assist ring scaled up inline (CSS is frozen). */
|
||||
function BigLoader({ title, children }) {
|
||||
return (
|
||||
<div style={{ display: 'grid', placeItems: 'center', padding: '64px 20px', textAlign: 'center' }}>
|
||||
<span
|
||||
className="ai-assist-spinner"
|
||||
role="status"
|
||||
aria-label={title}
|
||||
style={{ width: 72, height: 72, borderWidth: 5 }}
|
||||
/>
|
||||
<div className="fw-600" style={{ marginTop: 20, fontSize: 17 }}>{title}</div>
|
||||
{children && <p className="text-muted" style={{ marginTop: 6, maxWidth: 420 }}>{children}</p>}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
/** The JobCandidates MiniRing verbatim, fed by the deterministic job-match score. */
|
||||
function MatchRing({ score, size = 46 }) {
|
||||
if (score == null) return null
|
||||
const color = score >= 70 ? 'var(--success)' : score >= 40 ? 'var(--warning)' : 'var(--danger)'
|
||||
return (
|
||||
<div
|
||||
data-tip="Job match"
|
||||
style={{
|
||||
width: size, height: size, borderRadius: '50%', flexShrink: 0,
|
||||
display: 'grid', placeItems: 'center',
|
||||
background: `conic-gradient(${color} ${score}%, var(--bg-sunken) 0)`,
|
||||
}}
|
||||
>
|
||||
<div
|
||||
style={{
|
||||
width: size - 8, height: size - 8, borderRadius: '50%',
|
||||
background: 'var(--bg-elev)', display: 'grid', placeItems: 'center',
|
||||
fontWeight: 800, fontSize: 13.5, letterSpacing: '-.3px',
|
||||
color,
|
||||
}}
|
||||
>
|
||||
{score}
|
||||
</div>
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
function ProfileCard({ p, onView, onDismiss, dismissing }) {
|
||||
const crit = p.summary || p.headline || ''
|
||||
const shown = p.skills.slice(0, 5)
|
||||
const more = p.skills.length - shown.length
|
||||
return (
|
||||
<div className="card cand-card" onClick={() => onView(p)}>
|
||||
<div className="card-body">
|
||||
<div className="cand-head">
|
||||
<ProfileAvatar name={p.name} url={p.avatarUrl} />
|
||||
<div className="cand-id">
|
||||
<div className="cand-name">{p.name ?? 'Unknown'}</div>
|
||||
<div className="cand-role">{p.currentTitle ?? p.headline ?? '—'}</div>
|
||||
</div>
|
||||
<MatchRing score={p.matchScore} />
|
||||
</div>
|
||||
|
||||
{shown.length > 0 && (
|
||||
<div className="cand-skills">
|
||||
{shown.map((s) => <span className="cand-chip" key={s}>{s}</span>)}
|
||||
{more > 0 && <span className="cand-chip more">+{more} more</span>}
|
||||
</div>
|
||||
)}
|
||||
|
||||
<p className="cand-crit">{crit}</p>
|
||||
|
||||
<div className="cand-foot">
|
||||
<span className="cand-meta"><Icon name="map" /> {p.location ?? '—'}</span>
|
||||
<span className="cand-company">{p.currentCompany ?? ''}</span>
|
||||
<a
|
||||
className="act-btn"
|
||||
data-tip="Open LinkedIn profile"
|
||||
href={p.linkedinUrl}
|
||||
target="_blank"
|
||||
rel="noreferrer"
|
||||
onClick={(e) => e.stopPropagation()}
|
||||
>
|
||||
<Icon name="linkedin" />
|
||||
</a>
|
||||
<button
|
||||
className="act-btn"
|
||||
data-tip="View profile"
|
||||
onClick={(e) => { e.stopPropagation(); onView(p) }}
|
||||
>
|
||||
<Icon name="eye" />
|
||||
</button>
|
||||
<button
|
||||
className="act-btn"
|
||||
data-tip="Dismiss"
|
||||
disabled={dismissing}
|
||||
onClick={(e) => { e.stopPropagation(); onDismiss(p) }}
|
||||
>
|
||||
<Icon name="trash" />
|
||||
</button>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
/** Full LinkedIn profile: hero + about + skills + employment/education history. */
|
||||
function TalentProfileDetail({ profileId, onClose }) {
|
||||
const detailQuery = useQuery({
|
||||
queryKey: qk.talent.profile(profileId),
|
||||
queryFn: () => talentApi.getProfile(profileId),
|
||||
})
|
||||
const p = detailQuery.data?.data ? talentApi.toProfileDetailView(detailQuery.data.data) : null
|
||||
|
||||
return (
|
||||
<Modal
|
||||
title="Talent Profile"
|
||||
subtitle={p?.headline ?? undefined}
|
||||
size="modal-lg"
|
||||
onClose={onClose}
|
||||
footer={
|
||||
<>
|
||||
{p && (
|
||||
<a className="btn btn-primary" href={p.linkedinUrl} target="_blank" rel="noreferrer">
|
||||
<Icon name="linkedin" /> Open LinkedIn
|
||||
</a>
|
||||
)}
|
||||
<button className="btn" onClick={onClose}>Close</button>
|
||||
</>
|
||||
}
|
||||
>
|
||||
{detailQuery.isError ? (
|
||||
<EmptyState icon="alert" title="Could not load this profile">
|
||||
{friendlyAuthError(detailQuery.error, 'Please try again.')}
|
||||
</EmptyState>
|
||||
) : detailQuery.isPending ? (
|
||||
<BigLoader title="Loading profile…" />
|
||||
) : (
|
||||
<>
|
||||
<div className="profile-hero">
|
||||
<ProfileAvatar name={p.name} url={p.avatarUrl} />
|
||||
<div style={{ flex: 1 }}>
|
||||
<div className="ph-name">{p.name ?? 'Unknown'}</div>
|
||||
<div className="ph-role">
|
||||
{[p.currentTitle, p.currentCompany].filter(Boolean).join(' at ') || p.headline || '—'}
|
||||
</div>
|
||||
<div className="ph-tags">
|
||||
{p.location && <Badge className="b-plain b-indigo badge-plain">{p.location}</Badge>}
|
||||
<Badge className="b-gray">LinkedIn</Badge>
|
||||
{p.lastSeenAt && (
|
||||
<Badge className="b-plain b-indigo badge-plain">Found {fmtDate(p.lastSeenAt)}</Badge>
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{p.summary && (
|
||||
<>
|
||||
<div className="form-section-title">About</div>
|
||||
<p className="text-muted" style={{ whiteSpace: 'pre-line' }}>{p.summary}</p>
|
||||
</>
|
||||
)}
|
||||
|
||||
{p.skills.length > 0 && (
|
||||
<>
|
||||
<div className="form-section-title">Skills ({p.skills.length})</div>
|
||||
<div className="k-tags" style={{ marginBottom: 16 }}>
|
||||
{p.skills.map((s) => <span className="tag" key={s}>{s}</span>)}
|
||||
</div>
|
||||
</>
|
||||
)}
|
||||
|
||||
{p.experience.length > 0 && (
|
||||
<>
|
||||
<div className="form-section-title">Experience ({p.experience.length})</div>
|
||||
{p.experience.map((e, i) => (
|
||||
<div key={i} style={{ marginBottom: 14 }}>
|
||||
<div className="fw-600">
|
||||
{[e.title, e.company].filter(Boolean).join(' — ') || '—'}
|
||||
</div>
|
||||
<div className="text-muted text-sm">
|
||||
{[e.period, e.duration, e.employmentType, e.location].filter(Boolean).join(' · ')}
|
||||
</div>
|
||||
{e.description && (
|
||||
<p className="text-muted text-sm" style={{ marginTop: 4 }}>{e.description}</p>
|
||||
)}
|
||||
{e.skills.length > 0 && (
|
||||
<div className="cand-skills" style={{ marginTop: 6 }}>
|
||||
{e.skills.map((s) => <span className="cand-chip" key={s}>{s}</span>)}
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
))}
|
||||
</>
|
||||
)}
|
||||
|
||||
{p.education.length > 0 && (
|
||||
<>
|
||||
<div className="form-section-title">Education ({p.education.length})</div>
|
||||
{p.education.map((e, i) => (
|
||||
<div key={i} style={{ marginBottom: 12 }}>
|
||||
<div className="fw-600">{e.school ?? '—'}</div>
|
||||
<div className="text-muted text-sm">
|
||||
{[[e.degree, e.field].filter(Boolean).join(', '), e.period].filter(Boolean).join(' · ')}
|
||||
</div>
|
||||
</div>
|
||||
))}
|
||||
</>
|
||||
)}
|
||||
</>
|
||||
)}
|
||||
</Modal>
|
||||
)
|
||||
}
|
||||
|
||||
export default function Talent() {
|
||||
const { toast } = useToast()
|
||||
const qc = useQueryClient()
|
||||
|
||||
const [jobId, setJobId] = useState('')
|
||||
const [activeRunId, setActiveRunId] = useState(null)
|
||||
const [confirmOpen, setConfirmOpen] = useState(false)
|
||||
const [search, setSearch] = useState('')
|
||||
const [locationChoice, setLocationChoice] = useState('Pakistan')
|
||||
const [customLocation, setCustomLocation] = useState('')
|
||||
const [visibleCount, setVisibleCount] = useState(10)
|
||||
const [viewProfileId, setViewProfileId] = useState(null)
|
||||
|
||||
const jobsQuery = useQuery({ queryKey: qk.jobPosts.list(), queryFn: fetchJobs })
|
||||
const jobs = jobsQuery.data ?? []
|
||||
const selectedJob = jobs.find((j) => j.id === jobId)
|
||||
|
||||
const runsQuery = useQuery({
|
||||
queryKey: qk.talent.runs(jobId),
|
||||
queryFn: () => talentApi.listRuns(jobId),
|
||||
enabled: !!jobId,
|
||||
})
|
||||
const runs = useMemo(
|
||||
() => (Array.isArray(runsQuery.data?.data) ? runsQuery.data.data.map(talentApi.toRunView) : []),
|
||||
[runsQuery.data],
|
||||
)
|
||||
const latestRun = runs[0] ?? null
|
||||
|
||||
// Resume-after-reload: adopt the newest unfinished run as the poll target.
|
||||
useEffect(() => {
|
||||
if (!activeRunId && latestRun && !latestRun.isTerminal) setActiveRunId(latestRun.id)
|
||||
}, [activeRunId, latestRun])
|
||||
|
||||
const statusQuery = useQuery({
|
||||
queryKey: qk.talent.run(activeRunId),
|
||||
queryFn: () => talentApi.getRunStatus(activeRunId),
|
||||
enabled: !!activeRunId,
|
||||
refetchInterval: (query) => {
|
||||
const status = query.state.data?.data?.status
|
||||
return status && talentApi.isTerminalRun(status) ? false : 4000
|
||||
},
|
||||
})
|
||||
const activeRun = statusQuery.data?.data ? talentApi.toRunView(statusQuery.data.data) : null
|
||||
const runInFlight = !!activeRun && !activeRun.isTerminal
|
||||
|
||||
// Toast + refresh exactly once per run settling.
|
||||
const settledRef = useRef(null)
|
||||
useEffect(() => {
|
||||
if (!activeRun || !activeRun.isTerminal || settledRef.current === activeRun.id) return
|
||||
settledRef.current = activeRun.id
|
||||
qc.invalidateQueries({ queryKey: qk.talent.all() })
|
||||
setVisibleCount(10)
|
||||
if (activeRun.status === 'succeeded') {
|
||||
toast(`${activeRun.profilesFound} profile${activeRun.profilesFound === 1 ? '' : 's'} found on LinkedIn`, 'success')
|
||||
} else {
|
||||
toast(activeRun.error || `Talent search ${activeRun.status.replace('_', ' ')}`, 'error')
|
||||
}
|
||||
}, [activeRun, qc, toast])
|
||||
|
||||
const profilesQuery = useQuery({
|
||||
queryKey: qk.talent.profiles({ jobId }),
|
||||
queryFn: () => talentApi.listProfiles({ jobId }),
|
||||
enabled: !!jobId,
|
||||
})
|
||||
const profiles = useMemo(
|
||||
() =>
|
||||
Array.isArray(profilesQuery.data?.data)
|
||||
? profilesQuery.data.data.map(talentApi.toProfileView)
|
||||
: [],
|
||||
[profilesQuery.data],
|
||||
)
|
||||
const visible = useMemo(() => {
|
||||
const q = search.trim().toLowerCase()
|
||||
if (!q) return profiles
|
||||
return profiles.filter((p) =>
|
||||
[p.name, p.headline, p.currentCompany, p.currentTitle, p.location]
|
||||
.some((f) => f && f.toLowerCase().includes(q)),
|
||||
)
|
||||
}, [profiles, search])
|
||||
|
||||
const effectiveLocation =
|
||||
locationChoice === CUSTOM_LOCATION ? customLocation.trim() : locationChoice
|
||||
const locationLabel =
|
||||
locationChoice === 'Anywhere'
|
||||
? 'anywhere (no location filter)'
|
||||
: `in ${effectiveLocation}`
|
||||
|
||||
const starting = useMutation({
|
||||
mutationFn: () => talentApi.startRun(jobId, { location: effectiveLocation }),
|
||||
onSuccess: (res) => {
|
||||
setConfirmOpen(false)
|
||||
const run = res?.data
|
||||
if (run?.id) {
|
||||
qc.setQueryData(qk.talent.run(run.id), res)
|
||||
setActiveRunId(run.id)
|
||||
}
|
||||
qc.invalidateQueries({ queryKey: qk.talent.runs(jobId) })
|
||||
toast('Talent search started', 'success')
|
||||
},
|
||||
onError: (err) => {
|
||||
setConfirmOpen(false)
|
||||
toast(friendlyAuthError(err, 'Could not start the talent search'), 'error')
|
||||
},
|
||||
})
|
||||
|
||||
const dismissing = useMutation({
|
||||
mutationFn: (profile) => talentApi.deleteProfile(profile.id),
|
||||
onSuccess: () => qc.invalidateQueries({ queryKey: qk.talent.profiles({ jobId }) }),
|
||||
onError: (err) => toast(friendlyAuthError(err, 'Could not dismiss the profile'), 'error'),
|
||||
})
|
||||
|
||||
const statusRun = runInFlight || !latestRun ? activeRun : latestRun
|
||||
const [badgeCls, badgeLabel] = statusRun ? (RUN_BADGE[statusRun.status] ?? ['b-gray', statusRun.status]) : []
|
||||
|
||||
return (
|
||||
<div className="page">
|
||||
<div className="page-head">
|
||||
<div>
|
||||
<h1 className="page-title">Talent</h1>
|
||||
<p className="page-sub">Source matching LinkedIn profiles for a job via Apify</p>
|
||||
</div>
|
||||
<div className="page-head-actions">
|
||||
<span className="integration-status pending"><span className="pulse" />LinkedIn Sourcing · Live</span>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div className="card mb-18">
|
||||
<div className="card-body">
|
||||
<div className="flex items-center gap-8" style={{ flexWrap: 'wrap' }}>
|
||||
<select
|
||||
className="select"
|
||||
style={{ flex: 1, minWidth: 180 }}
|
||||
aria-label="Source for job"
|
||||
value={jobId}
|
||||
onChange={(e) => {
|
||||
const nextId = e.target.value
|
||||
setJobId(nextId)
|
||||
setActiveRunId(null)
|
||||
setSearch('')
|
||||
setVisibleCount(10)
|
||||
const preset = defaultLocationFor(jobs.find((j) => j.id === nextId))
|
||||
setLocationChoice(preset.choice)
|
||||
setCustomLocation(preset.custom)
|
||||
}}
|
||||
>
|
||||
<option value="">Select a job post…</option>
|
||||
{jobs.map((j) => <option key={j.id} value={j.id}>{j.title}</option>)}
|
||||
</select>
|
||||
<select
|
||||
className="select"
|
||||
style={{ width: 220 }}
|
||||
value={locationChoice}
|
||||
onChange={(e) => setLocationChoice(e.target.value)}
|
||||
aria-label="Location"
|
||||
>
|
||||
{LOCATION_OPTIONS.map((o) => (
|
||||
<option key={o.value} value={o.value}>{o.label}</option>
|
||||
))}
|
||||
</select>
|
||||
{locationChoice === CUSTOM_LOCATION && (
|
||||
<input
|
||||
className="input"
|
||||
style={{ width: 200 }}
|
||||
placeholder="City or country…"
|
||||
value={customLocation}
|
||||
onChange={(e) => setCustomLocation(e.target.value)}
|
||||
/>
|
||||
)}
|
||||
<button
|
||||
className="btn btn-primary"
|
||||
disabled={!jobId || runInFlight || starting.isPending || !effectiveLocation}
|
||||
onClick={() => setConfirmOpen(true)}
|
||||
>
|
||||
<Icon name={runInFlight ? 'clock' : 'search'} />
|
||||
{runInFlight ? 'Sourcing…' : 'Find Talent'}
|
||||
</button>
|
||||
</div>
|
||||
{jobsQuery.isError && (
|
||||
<p className="text-muted text-sm" style={{ marginTop: 12 }}>
|
||||
{friendlyAuthError(jobsQuery.error, 'Could not load job posts')}
|
||||
</p>
|
||||
)}
|
||||
{jobId && statusRun && (
|
||||
<p className="text-muted text-sm flex items-center gap-8" style={{ marginTop: 12 }}>
|
||||
<Badge className={badgeCls}>{badgeLabel}</Badge>
|
||||
{statusRun.status === 'succeeded' && (
|
||||
<span>{statusRun.profilesFound} profile{statusRun.profilesFound === 1 ? '' : 's'} in the last run</span>
|
||||
)}
|
||||
{statusRun.error && <span>{statusRun.error}</span>}
|
||||
{statusRun.createdAt && <span>· {fmtDate(statusRun.createdAt)}</span>}
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{!jobId ? (
|
||||
<EmptyState icon="user-plus" title="Pick a job to source for">
|
||||
Sourced LinkedIn profiles are saved per job and kept across searches.
|
||||
</EmptyState>
|
||||
) : profilesQuery.isError ? (
|
||||
<EmptyState icon="alert" title="Could not load sourced profiles">
|
||||
{friendlyAuthError(profilesQuery.error, 'Please try again.')}
|
||||
</EmptyState>
|
||||
) : profilesQuery.isPending ? (
|
||||
<BigLoader title="Loading sourced profiles…" />
|
||||
) : profiles.length === 0 ? (
|
||||
runInFlight ? (
|
||||
<BigLoader title="Searching LinkedIn…">
|
||||
Scanning profiles matching this job's title, skills and experience.
|
||||
This usually takes a minute or two — results appear here automatically.
|
||||
</BigLoader>
|
||||
) : (
|
||||
<EmptyState icon="search" title="No profiles sourced yet">
|
||||
Run Find Talent to search LinkedIn for people matching this job.
|
||||
</EmptyState>
|
||||
)
|
||||
) : (
|
||||
<>
|
||||
<div className="flex items-center gap-8 mb-18">
|
||||
<input
|
||||
className="input"
|
||||
style={{ maxWidth: 320 }}
|
||||
placeholder="Filter by name, headline, company…"
|
||||
value={search}
|
||||
onChange={(e) => setSearch(e.target.value)}
|
||||
/>
|
||||
<span className="text-muted text-sm">
|
||||
{visible.length} of {profiles.length} profile{profiles.length === 1 ? '' : 's'}
|
||||
</span>
|
||||
</div>
|
||||
<div className="grid g-3">
|
||||
{visible.slice(0, visibleCount).map((p) => (
|
||||
<ProfileCard
|
||||
key={p.id}
|
||||
p={p}
|
||||
onView={(profile) => setViewProfileId(profile.id)}
|
||||
onDismiss={(profile) => dismissing.mutate(profile)}
|
||||
dismissing={dismissing.isPending}
|
||||
/>
|
||||
))}
|
||||
</div>
|
||||
<div className="flex items-center gap-8" style={{ justifyContent: 'center', marginTop: 18 }}>
|
||||
{visible.length > visibleCount ? (
|
||||
<button className="btn" onClick={() => setVisibleCount((n) => n + 10)}>
|
||||
<Icon name="chevron-down" />
|
||||
Show more ({visible.length - visibleCount} remaining)
|
||||
</button>
|
||||
) : (
|
||||
<button
|
||||
className="btn"
|
||||
disabled={runInFlight || starting.isPending}
|
||||
onClick={() => setConfirmOpen(true)}
|
||||
>
|
||||
<Icon name={runInFlight ? 'clock' : 'search'} />
|
||||
{runInFlight ? 'Sourcing…' : 'Search LinkedIn for more'}
|
||||
</button>
|
||||
)}
|
||||
</div>
|
||||
</>
|
||||
)}
|
||||
|
||||
{viewProfileId && (
|
||||
<TalentProfileDetail profileId={viewProfileId} onClose={() => setViewProfileId(null)} />
|
||||
)}
|
||||
|
||||
{confirmOpen && (
|
||||
<Modal
|
||||
title="Start LinkedIn talent search"
|
||||
subtitle={selectedJob?.title}
|
||||
onClose={() => setConfirmOpen(false)}
|
||||
footer={
|
||||
<>
|
||||
<button className="btn" onClick={() => setConfirmOpen(false)}>Cancel</button>
|
||||
<button
|
||||
className="btn btn-primary"
|
||||
disabled={starting.isPending}
|
||||
onClick={() => starting.mutate()}
|
||||
>
|
||||
{starting.isPending ? 'Starting…' : 'Start search'}
|
||||
</button>
|
||||
</>
|
||||
}
|
||||
>
|
||||
<p>
|
||||
This starts a <strong>paid</strong> Apify search of LinkedIn for people matching
|
||||
this job's title, technical requirements and experience level, {locationLabel} —
|
||||
up to 25 profiles per run (roughly $0.20). Repeating the same search continues
|
||||
deeper into the results, so each run surfaces new people; anyone already found
|
||||
is refreshed, not duplicated.
|
||||
</p>
|
||||
</Modal>
|
||||
)}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
|
@ -0,0 +1,108 @@
|
|||
"""Live smoke test for the request shape. Makes two real API calls.
|
||||
|
||||
Run this once before trusting the service against a new model or SDK version:
|
||||
|
||||
python scripts/smoke_structured_output.py
|
||||
|
||||
It proves the three things unit tests cannot:
|
||||
|
||||
1. The schema derived from ``ATSScore`` is accepted by structured outputs, and the
|
||||
configured model supports both it and the requested reasoning effort.
|
||||
2. ``output_parsed`` comes back as a valid ``ATSScore``.
|
||||
3. The shared job-description prefix is actually cached -- the second call reports
|
||||
``usage.input_tokens_details.cached_tokens > 0``.
|
||||
|
||||
Needs OPENAI_API_KEY in the environment or .env, and spends a few cents.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
# Running a script directly puts scripts/ on sys.path[0], not the repo root. This
|
||||
# environment has another project on the path via an editable-install .pth file, and
|
||||
# it also ships a top-level `app` package -- without this line `import app` silently
|
||||
# resolves to that one instead.
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
|
||||
from openai import AsyncOpenAI
|
||||
|
||||
from app.core.config import get_settings, supports_reasoning
|
||||
from app.core.logging import configure_logging
|
||||
from app.services.llm import OpenAIScorer
|
||||
|
||||
# OpenAI only caches prompts at or above 1024 tokens, so a short job description will
|
||||
# report zero cached tokens no matter how stable the prefix is. This one clears it.
|
||||
JOB_DESCRIPTION = (
|
||||
"Senior backend engineer.\n\n"
|
||||
"Required: Python 3.12, FastAPI, asyncio, Docker, PostgreSQL, REST API design, "
|
||||
"and demonstrated ownership of production services.\n"
|
||||
"Preferred: AWS, Kubernetes, Terraform, observability tooling.\n\n"
|
||||
) + ("Responsibilities include designing, shipping, and operating backend services. " * 200)
|
||||
|
||||
RESUME_A = (
|
||||
"Ada Lovelace\nBackend engineer, 6 years.\n"
|
||||
"Built FastAPI services on Python 3.12 with asyncio and PostgreSQL. "
|
||||
"Owned Docker-based deploys and on-call for a payments API."
|
||||
)
|
||||
RESUME_B = (
|
||||
"Grace Hopper\nData engineer, 3 years.\n"
|
||||
"Primarily ETL in Python with pandas and Airflow. Familiar with REST APIs. "
|
||||
"No production service ownership listed."
|
||||
)
|
||||
|
||||
|
||||
async def main() -> int:
|
||||
settings = get_settings()
|
||||
configure_logging(level=settings.log_level, fmt=settings.log_format)
|
||||
print(
|
||||
f"model={settings.openai_model} "
|
||||
f"effort={settings.openai_effort if supports_reasoning(settings.openai_model) else 'n/a'} "
|
||||
f"max_output_tokens={settings.openai_max_output_tokens}"
|
||||
)
|
||||
|
||||
client = AsyncOpenAI(
|
||||
api_key=settings.openai_api_key or None,
|
||||
timeout=settings.openai_timeout_seconds,
|
||||
max_retries=settings.openai_max_retries,
|
||||
)
|
||||
scorer = OpenAIScorer(
|
||||
client,
|
||||
model=settings.openai_model,
|
||||
max_output_tokens=settings.openai_max_output_tokens,
|
||||
effort=settings.openai_effort,
|
||||
enable_cache=settings.openai_enable_prompt_cache,
|
||||
)
|
||||
|
||||
try:
|
||||
# Sequential on purpose: a cache entry is only readable once the first
|
||||
# response exists, which is exactly what score_batch's priming step does.
|
||||
first = await scorer.score(JOB_DESCRIPTION, RESUME_A)
|
||||
print(
|
||||
f"call 1 ok: score={first.match_score} name={first.candidate_name!r} "
|
||||
f"title={first.job_title!r} years={first.years_experience} "
|
||||
f"critique={first.summary_critique!r}"
|
||||
)
|
||||
|
||||
second = await scorer.score(JOB_DESCRIPTION, RESUME_B)
|
||||
print(
|
||||
f"call 2 ok: score={second.match_score} name={second.candidate_name!r} "
|
||||
f"title={second.job_title!r} years={second.years_experience} "
|
||||
f"critique={second.summary_critique!r}"
|
||||
)
|
||||
finally:
|
||||
await client.close()
|
||||
|
||||
print(
|
||||
"\nSchema accepted and both responses parsed. "
|
||||
"Check the 'candidate_scored_upstream' log lines above: call 2 should show a "
|
||||
"non-zero cached_tokens. If it is zero, either the prompt is under the 1024-token "
|
||||
"caching minimum or the job-description prefix is not byte-stable across calls."
|
||||
)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(asyncio.run(main()))
|
||||
Loading…
Reference in New Issue