01 / The project
Why I built it
I already had a job-market data pipeline and a hand-tuned AI writing workflow. This app brings them together: find a posting, keep its description, review the relevant applicant evidence, and prepare a draft without re-explaining that background each time.
The public demo lets visitors search the real job inventory and open descriptions. A fictional profile and guided walkthrough show how the private applicant workflow fits together. Saved profiles, uploads and AI calls are protected by authentication and currently limited to one pilot account.
This is also a writing experiment. I test its drafts blind against my existing workflow, record the losses, and keep the applicant in charge of reviewing, editing and submitting their materials.
02 / Architecture
The system, from boards to drafts
The pipeline owns shared job inventory. The app owns private applicant records. A visitor can search real postings and explore a fictional profile; private drafting sits behind sign-in.
Shared inventory & public search
- Job boardsGreenhouse, Ashby, Lever, USAJobs
- Daily ingestionPython collection and validated records
- Postgres + dbtHistory, classification, search mart
- Web appPublic search through a read-only reader
A selected posting enters the applicant's workspace
Private applicant workflow
- WorkspaceProfile evidence and saved description
- AI draftCompany research and cited fact IDs
- Review checksBlocking and advisory warnings
- ApplicantReviews, edits and sends
03 / Data engineering
A daily pipeline with memory
A scheduled Python ingestion job collects thousands of company job boards across Greenhouse, Ashby and Lever, plus USAJobs. Repeat runs update the same posting using its source identity, rather than inserting duplicates. dbt turns those records into tested models, with rule-based title classification for role family and seniority, plus location parsing.
Versioned snapshots, called SCD2, retain posting changes. They measure observed posting duration, not whether a vacancy was filled. A failed collection must not look like every job closed. Daily builds run data tests; CI rebuilds against fictional fixtures. Raw data has a 90-day retention policy. Its CI check is a dry run and deletes nothing.
Read the pipeline write-up for the collection and posting-history decisions.
04 / Application
Public discovery, private evidence
The app runs on Next.js and Vercel. Public search uses fixed, parameterized queries through a separate read-only database login. Rate limiting protects public endpoints. Applicant records stay separate from the shared warehouse, and uploads go to private object storage.
Authenticated features currently admit one pilot account. Visitors can use public search or follow the fictional walkthrough. The portfolio itself is static Astro; it holds this explanation and makes no private app or model calls.
05 / AI drafting
Evidence travels with the draft
- Facts have identities.
- Confirmed profile facts supply applicant evidence, and each draft lists the fact IDs it used. Unreviewed résumé lines can support a preliminary draft and are labelled separately. Writing samples teach style, never qualifications.
- Keep useful drafts; show the risks.
- Review checks attach warnings for unsupported metrics, tools the applicant does not have, placeholders, and unstated commitments such as relocation or start dates. Blocking warnings stop “Ready to apply” until reviewed. These checks assist human review; they do not prove every sentence true.
- Research needs a retrieved source.
- Company research uses web search. A returned citation is accepted only when its URL matches a source the search tool retrieved, and job boards are excluded. The applicant can check that source.
- Reserve the cost before calling.
- Every model call has an explicit price and a reservation against a budget ledger. An unknown price keeps AI off. Uncertain usage retains the reservation.
- Ask for what only the applicant knows.
- A story bank asks at most one optional question per application and saves the answer in the applicant's own words. Incident selection considers prior use so letters can draw on different examples.
06 / Evaluation
Can it write as well as my workflow?
I already had a hand-tuned workflow using Claude skills for research, writing rules and drafts in my voice. The app has to earn its place against that baseline.
The workflow still writes better letters.
Each case freezes a posting, question and limits. Blind A/B packets randomize order with a recorded seed, normalize revealing formatting, and keep the answer key separate until after rating. Preference and a short reason are the main signal.
An exact sign test excludes ties; a Wilson 95% interval shows uncertainty in the win rate. These rounds have at most ten cases, so they are pilot signals. Repeated development cases guide changes; fresh held-out cases test whether those changes transfer.
Scores read first arm wins / comparison arm wins / ties. Rounds 3 and 4 compare a feature enabled against disabled, not the app against the workflow. Scroll the table on a small screen.
| Round / change | Comparison | Wins / losses / ties | Rater | What it showed |
|---|---|---|---|---|
| 01Stronger model; drafts kept with warnings | App vs. workflow10 cases | 1 / 8 / 1 | Me | App was worse. Two-sided p = 0.039; win rate 11% (95% interval: 2–43%). |
| 02Researched, source-checked company detail | App vs. workflowSame 8 cases | 4 / 4 / 0 | Me | Even split. Win rate 50% (95% interval: 22–78%). |
| 03My writing as voice examples | With vs. without examples8 cases | 4 / 4 / 0 | ChatGPT stand-in | No preference advantage. Voice examples stay off. |
| 04Story bank with applicant answers | With vs. without stories8 cases | 3 / 5 / 0 | ChatGPT stand-in | No significant advantage (p = 0.73). More useful for “why us” answers than letters. |
| RésuméTailoring an approved base résumé | App vs. my custom résumé3 cases | 1 / 0 / 2 | ChatGPT stand-in | Competitive in this small sample. |
| Held outFull system on fresh postings | App vs. workflow3 new cases | 0 / 2 / 1 | ChatGPT stand-in | The workflow still writes better letters. |
The first round's reasons were specific: generic letters, little personal connection and clunky phrasing. Research brought round 2 to an even split; four cases improved and one worsened (one-sided p ≈ 0.19). That is a useful direction, not proof of superiority.
Rounds 3 onward use ChatGPT as my stand-in, with the blind packets and writing rules. Those ratings are weaker evidence than my own. The stand-in caught swapped coverage figures in both arms and traced them to the source profile. Source documents were corrected. Draft review gained an unstated-commitment check, and the prompt distinguishes company sources from third-party reporting.
Judge agreement
1 of 10
A blind LLM judge agreed with my preference on only one packet, despite consistent choices when A/B order was reversed. It rewarded grounded claims; I wanted a letter that sounded like me. I use it for factual-risk checks, not draft ranking.
Evaluation model spend
About $5.10
Across drafts, research and judge runs, against a $10 cap. This is the recorded evaluation spend, not a promise about future running costs.
07 / Next steps
The next test is actual use
- Close the cover-letter gap against the existing workflow.
- Measure how much applicants change a draft before using it, alongside preference and factual checks.
- Learn from a small, invite-only beta.
Want to use this with your own profile? I'm inviting a few people to a private beta. Email me at wjcc91@gmail.com.