← All work

October 2026 · Personal project · Built solo

Job Search Workspace

A job search app and a writing experiment. Search my daily job-market pipeline, explore the applicant workflow, and see how its AI drafts hold up against my own writing.

01 / The project

Why I built it

I already had a job-market data pipeline and a hand-tuned AI writing workflow. This app brings them together: find a posting, keep its description, review the relevant applicant evidence, and prepare a draft without re-explaining that background each time.

The public demo lets visitors search the real job inventory and open descriptions. A fictional profile and guided walkthrough show how the private applicant workflow fits together. Saved profiles, uploads and AI calls are protected by authentication and currently limited to one pilot account.

This is also a writing experiment. I test its drafts blind against my existing workflow, record the losses, and keep the applicant in charge of reviewing, editing and submitting their materials.

02 / Architecture

The system, from boards to drafts

The pipeline owns shared job inventory. The app owns private applicant records. A visitor can search real postings and explore a fictional profile; private drafting sits behind sign-in.

Shared inventory & public search

  1. Job boardsGreenhouse, Ashby, Lever, USAJobs
  2. Daily ingestionPython collection and validated records
  3. Postgres + dbtHistory, classification, search mart
  4. Web appPublic search through a read-only reader

A selected posting enters the applicant's workspace

Private applicant workflow

  1. WorkspaceProfile evidence and saved description
  2. AI draftCompany research and cited fact IDs
  3. Review checksBlocking and advisory warnings
  4. ApplicantReviews, edits and sends
The app prepares materials. The applicant makes the final decisions and submits them.

03 / Data engineering

A daily pipeline with memory

A scheduled Python ingestion job collects thousands of company job boards across Greenhouse, Ashby and Lever, plus USAJobs. Repeat runs update the same posting using its source identity, rather than inserting duplicates. dbt turns those records into tested models, with rule-based title classification for role family and seniority, plus location parsing.

Versioned snapshots, called SCD2, retain posting changes. They measure observed posting duration, not whether a vacancy was filled. A failed collection must not look like every job closed. Daily builds run data tests; CI rebuilds against fictional fixtures. Raw data has a 90-day retention policy. Its CI check is a dry run and deletes nothing.

Read the pipeline write-up for the collection and posting-history decisions.

04 / Application

Public discovery, private evidence

The app runs on Next.js and Vercel. Public search uses fixed, parameterized queries through a separate read-only database login. Rate limiting protects public endpoints. Applicant records stay separate from the shared warehouse, and uploads go to private object storage.

Authenticated features currently admit one pilot account. Visitors can use public search or follow the fictional walkthrough. The portfolio itself is static Astro; it holds this explanation and makes no private app or model calls.

05 / AI drafting

Evidence travels with the draft

Facts have identities.
Confirmed profile facts supply applicant evidence, and each draft lists the fact IDs it used. Unreviewed résumé lines can support a preliminary draft and are labelled separately. Writing samples teach style, never qualifications.
Keep useful drafts; show the risks.
Review checks attach warnings for unsupported metrics, tools the applicant does not have, placeholders, and unstated commitments such as relocation or start dates. Blocking warnings stop “Ready to apply” until reviewed. These checks assist human review; they do not prove every sentence true.
Research needs a retrieved source.
Company research uses web search. A returned citation is accepted only when its URL matches a source the search tool retrieved, and job boards are excluded. The applicant can check that source.
Reserve the cost before calling.
Every model call has an explicit price and a reservation against a budget ledger. An unknown price keeps AI off. Uncertain usage retains the reservation.
Ask for what only the applicant knows.
A story bank asks at most one optional question per application and saves the answer in the applicant's own words. Incident selection considers prior use so letters can draw on different examples.

06 / Evaluation

Can it write as well as my workflow?

I already had a hand-tuned workflow using Claude skills for research, writing rules and drafts in my voice. The app has to earn its place against that baseline.

The workflow still writes better letters.

Each case freezes a posting, question and limits. Blind A/B packets randomize order with a recorded seed, normalize revealing formatting, and keep the answer key separate until after rating. Preference and a short reason are the main signal.

An exact sign test excludes ties; a Wilson 95% interval shows uncertainty in the win rate. These rounds have at most ten cases, so they are pilot signals. Repeated development cases guide changes; fresh held-out cases test whether those changes transfer.

Scores read first arm wins / comparison arm wins / ties. Rounds 3 and 4 compare a feature enabled against disabled, not the app against the workflow. Scroll the table on a small screen.

All evaluation rounds, recorded October 9, 2026
Round / changeComparisonWins / losses / tiesRaterWhat it showed
01Stronger model; drafts kept with warningsApp vs. workflow10 cases1 / 8 / 1MeApp was worse. Two-sided p = 0.039; win rate 11% (95% interval: 2–43%).
02Researched, source-checked company detailApp vs. workflowSame 8 cases4 / 4 / 0MeEven split. Win rate 50% (95% interval: 22–78%).
03My writing as voice examplesWith vs. without examples8 cases4 / 4 / 0ChatGPT stand-inNo preference advantage. Voice examples stay off.
04Story bank with applicant answersWith vs. without stories8 cases3 / 5 / 0ChatGPT stand-inNo significant advantage (p = 0.73). More useful for “why us” answers than letters.
RésuméTailoring an approved base résuméApp vs. my custom résumé3 cases1 / 0 / 2ChatGPT stand-inCompetitive in this small sample.
Held outFull system on fresh postingsApp vs. workflow3 new cases0 / 2 / 1ChatGPT stand-inThe workflow still writes better letters.

The first round's reasons were specific: generic letters, little personal connection and clunky phrasing. Research brought round 2 to an even split; four cases improved and one worsened (one-sided p ≈ 0.19). That is a useful direction, not proof of superiority.

Rounds 3 onward use ChatGPT as my stand-in, with the blind packets and writing rules. Those ratings are weaker evidence than my own. The stand-in caught swapped coverage figures in both arms and traced them to the source profile. Source documents were corrected. Draft review gained an unstated-commitment check, and the prompt distinguishes company sources from third-party reporting.

Judge agreement

1 of 10

A blind LLM judge agreed with my preference on only one packet, despite consistent choices when A/B order was reversed. It rewarded grounded claims; I wanted a letter that sounded like me. I use it for factual-risk checks, not draft ranking.

Evaluation model spend

About $5.10

Across drafts, research and judge runs, against a $10 cap. This is the recorded evaluation spend, not a promise about future running costs.

07 / Next steps

The next test is actual use

  • Close the cover-letter gap against the existing workflow.
  • Measure how much applicants change a draft before using it, alongside preference and factual checks.
  • Learn from a small, invite-only beta.

Want to use this with your own profile? I'm inviting a few people to a private beta. Email me at wjcc91@gmail.com.