Commit Graph
86 Commits
Author SHA1 Message Date
Jakob Stender Guldberg c93609cd22 fix(gmail-sync,outcome): keep free-form tracker notes free of CSV-breaking characters (#454) (#455)
* fix(gmail-sync): strip CSV-breaking characters from the email subject

Step 7a interpolated the raw subject line of a received email into the
`notes` column of job_search_tracker.csv. No writer here emits a quoted
tracker field, so an unescaped comma splits the row - for csv.DictReader
just as much as for a naive split, which matters because
tools/rank_state.py is the repo's only machine reader and uses exactly
that. `notes` is column 10 of 14, so a subject as ordinary as
"Re: Your application, Data Scientist" shifted cv_file,
cover_letter_file and source a column left. A line break is worse: it
ends the row and starts a second one.

The rule now sits on the append instruction itself rather than in a
general note a writer can miss. The subject survives verbatim in the
archive's outcome.md, which is Markdown and carries no such constraint.

/outcome Step 4 (outcome.md:195) is also free-form and has the same
exposure, but its text is model-authored in a turn the user is watching
rather than copied from third-party mail unattended. Left out
deliberately, to be filed separately.

* fix(outcome): keep the Step 4 tracker note free of CSV-breaking characters

/outcome Step 4 appended "a short dated note" to `notes` with no
constraint on its content, the same exposure /gmail-sync Step 7a had:
nothing quotes a tracker field, so `rejected, no feedback given` shifts
cv_file, cover_letter_file and source a column left under
csv.DictReader, and a line break ends the row. The append instruction
now requires a note with no commas, double quotes or line breaks.

Folded in at the maintainer's request on #455 so one entry and one rule
cover both free-form writers. The CSV-safety tests move out of
test_gmail_sync_command.py into test_tracker_notes_csv_safe.py, where a
CASES table pins the rule on each writer's append line.
2026-09-16 06:45:55 +02:00
Souptik Chakraborty 1b65f7198a fix(verify-layout): name both causes of a broken bounding-box extractor (#451) (#465)
verify_layout.py's skipped: message blamed only the xpdf-based pdftotext that
Git for Windows puts ahead of Poppler in PATH. A real Poppler can abort too:
26.0x before 26.05 crashes -bbox/-bbox-layout/-htmlmeta on a PDF whose Info
dictionary carries an empty string in any field, and hyperref writes exactly
that for every field it does not set. A lualatex/pdflatex document built with
hyperref and no \hypersetup{pdftitle=...} - an ordinary /add-template CV
template, not a malformed one - hits this with a working Poppler installed,
and the old message sent the reader to check their PATH when nothing was
wrong with it.

Named both causes in the raised message, the code comment above it, and the
module docstring. Behavior is unchanged: either cause still degrades to
skipped: exit 2, not exit 1, since a broken extractor is still not a broken
document.

Added test_poppler_abort_on_empty_info_string_names_that_cause_too, mocking
subprocess.run to raise CalledProcessError the way a real Poppler 26.0x abort
does (exit 1, a libc++abi out_of_range trace on stderr) rather than the
xpdf case's exit 99 - the two failures are the same exception type with
different exit codes and stderr text, so the message has to actually
distinguish them, not just catch the one class. Negative control: this test
fails against the pre-fix message with "'Poppler aborted' not found in
'...a pdftotext without -bbox is usually the xpdf build...'".

Thanks to @main-sounds-audio for the report and the isolated repro (which
field, which Poppler modes, five runs of five) that pinned this to Poppler's
own upstream regression (issue #1699, fixed in 26.05.0) rather than an
extraction-library swap.

Fixes #451
2026-09-16 06:43:20 +02:00
Oscar Madera 88d1b61047 feat(apply): verify posting source host against installed portals and known ATS apexes (#431) (#467)
- Add source host verification rule to apply.md Step 1 before drafting
- Check posting URL host against installed portals and six official ATS apexes
- Enforce fail-closed look-alike parsing (prefix, suffix, userinfo spoofing)
- Plainly flag unverified third-party hosts in evaluation output
- Add tests/test_apply_host_check.py and update CHANGELOG.md
2026-09-16 06:43:16 +02:00
NathanandNathan No-ot 9484a61831 fix(ci): skip the pristine-template guard in test_setup_command on forks (#463)
TemplatesStillCarryThePlaceholders asserts 05-cv-templates.md and 06-cover-letter-templates.md still carry their [FIRST_NAME]/[YOUR_*] tokens. /setup's documented Step 3.5/3.6 personalisation replaces exactly those tokens, so python-tests failed permanently on any personalized fork. Same @unittest.skipIf on GITHUB_REPOSITORY (defaulting to upstream when unset) as test_placeholder_integrity.py received in #407; that fix landed three days before this guard, which did not pick up the pattern.

Co-authored-by: Nathan No-ot <244263078+Nnoot02@users.noreply.github.com>
2026-09-14 18:35:23 +02:00
Ayobami Adegoke 73d52e0991 fix(verify_pdf): fold LaTeX's typographic substitutions before --contains; guard T1 fontenc for pdflatex (#385, #384) (#458)
`normalize_text()` folded whitespace only, so `--contains` compared what a
user types against what LaTeX renders. The stock CV compiled with the
documented lualatex command turns `'` into U+2019 and `--` into U+2013, so
`--contains "Master's degree"` and `--contains "2016-2024"` both reported
the keyword missing from a document that plainly contains it, through both
extractors. The documented remedy for a missing keyword is to add it, which
is the one thing the ATS section forbids.

Fold both sides at comparison time: NFC, then curly apostrophes and quotes
to ASCII, en/em dashes to `-`, no-break space to space. `--dump-text` still
writes the raw layer - that is what an ATS parses, and the date-range rule
in 05-cv-templates.md needs the raw en-dash visible there.

Separately, pdflatex without T1 font encoding stores accents decomposed
(`e` + U+0300). NFC repairs the pdftotext side of that, but pypdf reads the
same layer as `Z¨ urich` with a spacing accent, which no fold recovers.
moderncv 2.5 loads T1 itself under pdflatex; the apt-packaged 2.3.1 does
not - reproduced by compiling the template against moderncv v2.3.1 with
pdflatex (before: U+0308/U+0300 in pdftotext, `Z¨ urich` in pypdf; after:
U+00FC/U+00E8 in both). The template and the guide's preamble gain
`\ifpdftex\usepackage[T1]{fontenc}\fi`; the lualatex text layer is
byte-identical before and after.

Tests: ten new cases in test_verify_pdf.py (the fold-through and
normalize_text ones fail on the whitespace-only code) and a
test_latex_guidance.py guard that the fontenc line exists and stays inside
the pdflatex branch. framework_version 1.4.3 -> 1.4.4 on 05-cv-templates.md.

Reported and diagnosed by 9scorp4 in Discussions #385 and #384.
2026-09-14 18:30:50 +02:00
Jakob Stender Guldberg c776e3f2b6 fix(outcome,interview): glob the full <company>_<role> stem in the CV fallback (#444)
When a tracker row's cv_file/cover_letter_file columns are empty, /outcome
and /interview fell back to a company-prefix glob (cv/main_<company>*.tex).
/apply names drafts main_<company>_<role><CV_EXT>, so two roles at one
company both match that glob. /outcome copied whichever the filesystem
returned first into the archive as cv_draft.tex - the file whose stated
purpose is to record what was actually submitted - and its own "leave an
existing archived file" rule then made the wrong copy permanent.

Both fallbacks now glob cv/main_<company>_<role>.* and
cover_letters/cover_<company>_<role>.*, deriving the stem by the Subfolder
naming rule in documents/README.md rather than restating it, and skip with
a note instead of widening the search. Dropping the hardcoded .tex also
makes a template registered by /add-template findable.

The dot before the extension wildcard matters: a bare trailing * also
absorbs a longer role, so ML Engineer and ML Engineer II at one company
would collide the same way the company-prefix glob did.
2026-09-10 20:32:37 +02:00
cbd8a991ab feat(layout): measure compiled PDF layout instead of eyeballing it (#378)
* feat(layout): measure compiled PDF layout instead of eyeballing it

/apply Step 5b asks for layout properties and executes none of them: they are
checked by reading the rendered page, which is exactly how they get missed.

The failure that motivates this is silent under every existing check. A moderncv
\cventry renders as a tabular, so an entry is one unbreakable block; when it does
not fit in the space left, the whole entry moves to the next page and leaves a
hole behind. The document still compiles, still reports the expected page count,
and still passes tools/verify_pdf.py. Observed in the wild at 273pt, roughly 19
blank lines, mid-page, on a CV whose visual read looked fine.

tools/verify_layout.py reports per page where the text starts and stops, bottom
whitespace as a share of page height, and the largest gap between lines, then
exits 1 on a hole over 100pt, a non-final page ending more than 25% early, body
text colliding with the page-number footer, a final page more than 35% empty, or
an entry header or section heading stranded at a page break.

Page count is deliberately not checked here - verify_pdf.py --pages already does
that, and two implementations of one rule drift. Geometry comes from Poppler
pdftotext -bbox, already a dependency; a missing Poppler exits 2 with "skipped:"
rather than failing the run.

Tests build synthetic Page/Line geometry, so the suite needs neither Poppler nor
a LaTeX toolchain and runs on the existing 3.10-3.14 matrix.

* fix(layout): survive Windows encoding and a pdftotext without -bbox

Review found two failures on the repo's primary platform:

subprocess.run(..., text=True) decoded pdftotext's UTF-8 output with the
Windows ANSI codepage and crashed on the stock cv/main_example.pdf. It now
passes encoding="utf-8" with errors="replace" at the call site, matching the
fix verify_pdf.py already carries from #369.

Git for Windows ships an xpdf-based pdftotext that shadows Poppler in a
default PATH and has no -bbox flag; it exited 99, the CalledProcessError
escaped, and the run ended in exit 1 - indistinguishable from a real layout
problem, which would send /apply chasing a phantom hole. That now routes to
the existing "skipped:" exit 2 path with a message naming the likely cause,
covered by three tests on the extractor-failure path.

Also from review:

- .claude/settings.json and security_guards.py gain the verify_layout.py
  permission entries. apply.md Step 5b runs the tool on every /apply, so
  without them every run prompts.
- The docstring and CHANGELOG no longer call Poppler a dependency
  verify_pdf.py relies on. Since #369 verify_pdf prefers pypdf and Poppler is
  the fallback; word bboxes have no pypdf equivalent, so this is the one step
  that still wants it, and that is now what the text says.
- Docstring and apply.md state that the thresholds are calibrated for the
  stock moderncv and cover.cls geometry, and that the shipped example CV
  fails the thin-final-page rule by design.
- largest_gap documents that it measures top-to-top, so a tall line inflates
  the gap by its own height - over-detection, the safe direction.
- The test module imports via tools.verify_layout like test_verify_pdf.py
  instead of sys.path.insert.

* fix(verify_layout): quote only the first stderr line in the skip message (#378)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013fqqLgQSnwgWkv98twQhHi

---------

Co-authored-by: Mads Lorentzen <madslorentzen_17@hotmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-09 19:26:49 +02:00
8c81edc330 fix(scrape): make seen_jobs.json keys a pure function of the posting (#441)
* fix(scraper): make seen_jobs.json keys a pure function of the posting, and clean up the drift

The dedup key was prose only ("<url_or_company_title_key>"), so different
/scrape runs slugified company+title differently and the state file
accumulated two failures: keys carrying "&", "/", "," and ":" that break
the archive-folder path /apply and /outcome derive from company+role
(documents/README.md's subfolder rule exists because of exactly this),
and the same posting stored twice under two different truncations of a
long title (two Deloitte entries, one job, one URL).

tools/job_key.py makes the key a pure, deterministic function of
company+title+url: a strict allowlist slug, length-capped with a hash of
the full slug so truncation never collides across runs, and a fallback
to the portal's numeric job id when a non-Latin title slugifies to
nothing (a real prior entry, "securion_", would have collided with
every future non-Latin posting from that company).

--audit finds both failure classes in an existing seen_jobs.json without
guessing at a fix: malformed keys (real damage), a legacy three-part
company_title_location shape (harmless but not what the current rule
produces, so it silently re-duplicates on the next scrape), and
duplicate URLs. Ran it against this workspace's file and re-keyed the 15
entries it found - 7 malformed, 8 legacy-shape - verified byte-for-byte
against the pre-cleanup copy that no entry's data changed, only its key.

tests/test_job_key.py (16 tests) covers the slugify rules, the
truncation-hash behavior, both non-Latin fallback paths, and the audit
CLI's exit codes.

* fix(scrape): call the key helper from Step 4 instead of slugifying ad hoc

The helper added in the previous commit is only load-bearing if the spec
calls it. Step 4 described the key as prose ("<url_or_company_title_key>"),
which is what let each run slugify its own way. Step 4 now names the
command, and the schema shows the key's provenance.

Renumbers the trailing list item; no other behaviour in the step changes.

* docs(changelog): record the job-key rule under Unreleased

* fix(scrape): preserve dedup continuity across key rule

* changelog: note that existing seen_jobs.json entries need no migration (#441)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013fqqLgQSnwgWkv98twQhHi

---------

Co-authored-by: nox <nox@Mac.home>
Co-authored-by: Mads Lorentzen <madslorentzen_17@hotmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-08 16:25:15 +02:00
Oscar Madera 6b07b13bd2 feat(commands): add GitHub repository project extraction to /expand (#437) 2026-09-07 18:43:59 +02:00
Oscar Madera 6176e6aaca feat(outcome): add stale sweep branch for batch-resolving quiet applications (#439)
* feat(outcome): add stale sweep branch for batch-resolving quiet applications

- Add /outcome stale [N] and /outcome sweep [N] to Step 0
- Offer stale sweep in Step 1.3 when rows are quiet 60+ days
- Add Step 2c Stale Sweep Branch with interactive all/select/skip confirmation
- Batch-resolve qualifying open applications to canonical no_response
- Update tracker status and append dated notes; update archive outcome.md
- Add Rule 9 to Important Rules
- Add tests/test_outcome_stale.py and update CHANGELOG.md

* test(outcome): drop auxiliary candidate filter from spec test per review
2026-09-07 18:41:16 +02:00
Mads LorentzenandClaude Fable 5.1 1a116b3c64 chore(release): cut v1.7.1 (#434)
Rolls [Unreleased] into [1.7.1] - 2026-09-06 and moves the compare links.
Keeps an empty [Unreleased] heading above the release, per Keep a Changelog,
and teaches the CHANGELOG structure guard that an absent [Unreleased]
section is nothing to check rather than an error.


Claude-Session: https://claude.ai/code/session_013fqqLgQSnwgWkv98twQhHi

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 20:19:17 +02:00
Mads LorentzenandClaude Fable 5.1 e6f6f4e322 fix(setup): fill the CV and cover-letter template contact blocks; guard CHANGELOG structure (#433)
/setup Step 3 personalised cv/main_example.tex but never the LaTeX contact
blocks embedded in 05-cv-templates.md and 06-cover-letter-templates.md, the
two files /apply actually compiles from; 06 was not a Step 3 target at all.
Step 3.5 now names the 05 contact tokens, a new Step 3.6 covers the 06
contact line and signature, the completion summary lists 06, and /reset
restores both blocks instead of listing 06 as framework-only (the existing
/reset coverage test forced that half).

tests/test_changelog_structure.py checks [Unreleased] on every PR for
duplicate headings, unknown headings, orphan entries and conflict markers -
the #425 duplicate-heading shape that was fixed by hand at merge time.


Claude-Session: https://claude.ai/code/session_013fqqLgQSnwgWkv98twQhHi

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 11:51:51 +02:00
3bf41149e0 fix(rank): move seen_jobs.json read/write off /rank's per-run cost (#395) (#425)
* fix(rank): move seen_jobs.json read/write off the state file's own critical path (#395)

/rank's Step 1 read the whole of seen_jobs.json into the conversation to
select candidates by eye, and Step 4 emitted it back to record scores.
That cost is paid on every run regardless of how many jobs are scored,
and it grows for the life of the workspace, since the file is
append-only and most stored entries are `skipped`.

tools/rank_state.py moves that traffic into code:

- `candidates` selects entries per Step 1's existing rules (status
  filter, tracker exclusion, focus filter, `--limit`/`--all` from #424)
  and projects only the fields a scoring agent needs.
- `sweep` runs rule 6's expiry pass over entries the run did not
  re-score - a stored-date comparison, no fetch, no agent - preserving
  its defensive parsing of non-ISO deadlines and its `--all`
  reversibility.
- `apply` writes scoring results back atomically and prints the
  ranked/vetoed/expired rows Step 5's report is built from, preserving
  Step 4's existing write-back rules exactly: the `location` ->
  `location_verdict` legacy migration, the deadline
  null-is-not-a-correction rule, verbatim strengths/gaps persistence,
  and idempotent re-scoring.

Step 1, Step 3's rule 6, and Step 4 now route through the tool instead
of describing a manual read/write. Nothing about scoring policy changes
- no new status, no new persisted field, no change to what counts as a
veto. The tracker stays read-only and every write is atomic (temp file
+ rename).

tests/test_rank_state.py (25 tests) covers the three subcommands
directly. The new spec-guard class in test_rank_command.py derives the
fields Step 4 must preserve from Step 2's own JSON schema block rather
than retyping them as a second list, so a future edit to that contract
is what the test reads instead of something that can drift from it.

* fix(rank): add CHANGELOG entry and remove the undefined $SCRATCHPAD reference

Two mechanical fixes from review:

- Step 4 named the results hand-off file via $SCRATCHPAD, a variable
  nothing in the repo defines - a reader following the spec literally
  has no path to substitute. Named the location in prose instead (a
  temporary file outside the repo tree, never committed) and replaced
  the shell-variable-looking path in the example command with an
  explicit placeholder.
- Added the [Unreleased] entry this change was missing; the one
  already in the diff belongs to #424.

* changelog: fold the #395 entry into the existing Fixed section

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013fqqLgQSnwgWkv98twQhHi

---------

Co-authored-by: nox <nox@Mac.home>
Co-authored-by: Mads Lorentzen <madslorentzen17@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 10:49:20 +02:00
Instinct fd89eac178 fix(rank): bound scoring batches with --limit (#395) (#424) 2026-09-03 19:44:17 +02:00
6f0178a8a1 fix(ci): skip placeholder-integrity tests on forks in python-tests (#407)
Add @unittest.skipIf on GITHUB_REPOSITORY to TestCvSentinelsAreDataLocated
and TestProfileSentinelIsDataLocated so python-tests matches the upstream-only
placeholder-integrity job. Default unset GITHUB_REPOSITORY to upstream so local
pristine-template runs still execute the guards.

Fixes MadsLorentzen/ai-job-search#405

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: shahidbeig-a11y <shahidbeig-a11y@users.noreply.github.com>
2026-09-03 19:14:01 +02:00
OluwaJomilojuandClaude Sonnet 5 7f709eda57 fix(salary): require corroboration before accepting a header row (#415)
* fix(salary): require corroboration before accepting a header row

Header-row detection accepted the first row (of the first 10) where any
cell merely contained a company-pattern word - no check that the row
actually looked like a header. A source-citation row above the real
header table (standard in real Danish union/statistics exports, e.g.
"Kilde: ... opdelt efter arbejdsgiver ...") tripped it purely because
"arbejdsgiver" appeared in prose. The real header row then parsed as
data (its "Firma" cell became a bogus company), and every genuine
company silently lost all its salary data - exit 0, no warning.

A candidate row is now only accepted when a second cell also matches a
city/count/index pattern, and a sheet that ends up with zero detected
salary columns prints a warning instead of reporting success silently.

Fixes #414.

* fix(salary): require cross-cell corroboration, fall back for untyped columns

Two edge cases found in review of the corroboration fix:

- Same-cell corroboration wasn't enough: a citation sentence can pack a
  count-pattern word into the same sentence as the company-pattern one
  ("...opdelt efter arbejdsgiver, antal svar 1234"), which still passed
  the gate. Corroboration must now come from a different cell.

- The corroboration requirement itself broke sheets whose only real
  header has purely untyped salary columns (e.g. "Base pay 2025" /
  "Bonus 2025" - neither matches a known city/count/index pattern), so
  header detection found nothing at all. Falls back to the original
  any-cell-mentions-company rule when the strict pass finds no row in
  the first 10.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-03 08:19:08 +02:00
Ayobami Adegoke 284dc4c2d0 feat(rank): flag stale postings from the stored posted_date (#390) (#406)
Step 3 gains rule 7: a posting whose stored posted_date is more than 30
days old at rank time carries a visible staleness marker with its age
spelled out alongside the score - FLAG treatment like location and
language, never an exclusion. No posted_date or null means no flag and no
guess (never inferred from first_seen), and rule 6's defensive-parse rule
applies wherever the stored value is compared. Age is re-derived each run
and never persisted. Four new spec pins in test_rank_command.py, each
verified to fail against the rule-less spec.
2026-09-02 20:02:38 +02:00
soumyadip sarkarandClaude Sonnet 5 9833a5dcb7 fix(salary): treat null metadata/categories as absent instead of crashing (#413)
--validate treats an explicit "metadata": null / "categories": null the same
as an omitted key ("...must be an object when provided", None is skipped), but
format_entry read both through dict.get(key, {}), which only substitutes the
default for an *absent* key - a present-but-null value passed through. The
renderer then hit None.get("index_label", ...) (AttributeError) or, via the
numeric-field fallback, None[key] = value (TypeError), so a hand-maintained
salary_data.json using null for "no value" died with an uncaught traceback
right after printing "Found 1 match(es)".

format_entry now coerces both to {} up front, honouring the validator's
existing "when provided" contract at the single consumer that broke it.

Tests (all verified to fail on the unfixed renderer):
- two unit cases calling format_entry with null metadata / null categories
- two end-to-end cases running main() --validate (blesses the file) then the
  lookup path (renders it), one per null shape

Plus an [Unreleased] CHANGELOG entry.

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 21:35:59 +02:00
Jakob Stender Guldberg 4c38f7ce4c fix(security): move the interview protection note to the rule that provides it (#337)
The two-line comment above `documents/interview/**` says interview prep and
experience records live there. Nothing has ever written to that directory:
/interview saves its pack to
documents/applications/<company>_<role>/interview_prep_<stage>.md, covered by
the documents/applications/** rule. `git grep documents/interview` returns
only the two declarations of the rule itself (.gitignore and
REQUIRED_IGNORE_RULES), `git log --all -- 'documents/interview*'` is empty,
and documents/README.md documents the applications path outright.

Nothing leaks - the comment is the defect, and it is the misleading kind. It
is the one dedicated, well-argued line about interview material in the
personal-data block, so an auditor checking that the framework's most
sensitive artifact is covered reads it and stops, at the only path in the
block with no writer.

The comment's description of what needs protecting was always right; only its
location was wrong. It now sits above documents/applications/**, the rule that
actually provides that protection, so a reader auditing the block finds the
reasoning attached to the rule doing the work. documents/interview/** stays -
REQUIRED_IGNORE_RULES pins it, so dropping it from .gitignore alone turns CI
red, and it is harmless defence in depth - relabelled in both files as
belt-and-braces rather than the primary guard.

The new check-ignore case in GitignorePatternBehaviorTests derives the
prep-pack path from /interview's own spec instead of hardcoding it. That
distinction is the whole value of the test: a hardcoded path pins only that
documents/applications/** still matches that shape, which security_guards.py
already catches first, and stays green if /interview moves its output -
leaving the corrected comment stale exactly the way this issue found it.
Since #329 the spec states the location in two pieces - Step 1 derives the
archive folder, Step 3 names interview_prep_<stage>.md - so the test pins both
fragments separately and composes the concrete path from them. Mutation-
verified on each half: repointing the folder at documents/prep_packs/, and
renaming the file, both fail this test while `python3 tools/security_guards.py`
still reports OK.

The class's temp-repo setup moved to setUp for the second case.

ayobamiseun reviewed the pre-rebase branch and called all three rebase hazards
in advance: the split literal, the released CHANGELOG context, and the setUp
re-merge. Reached independently here during the rebase; the review was posted
first.
2026-08-31 17:49:22 +02:00
Sandun Wijerathne ea2f25b39c fix(scrape): persist each posting's publication date in seen_jobs.json (#390) (#391)
* fix(scrape): persist each posting's publication date in seen_jobs.json (#390)

Step 2's contract guarantees a `date` on every portal CLI's search output and
CI enforces it in test_scrape_contract.py; Step 3 uses that date to scope a run
to the last 14 days. Step 4's storage schema then dropped it, so a posting's age
was unrecoverable the moment the run ended - `first_seen` records when the
scraper saw an entry, not when the employer posted it. /rank reads the stored
entry rather than the run, so it had no age signal to weigh.

A freehire-search posting dated 2024-05-13 was scraped 27 months later and
ranked Strong Fit at position 1 of 133. The scoring note observed the listing
"may be long stale" in prose nothing reads, and an /apply run drafted a tailored
CV and cover letter against it.

The schema gains `posted_date` (null when the portal returned no date, never
inferred or backfilled), documented alongside `deadline` with the same
never-backfill rule. Three new cases, each verified to fail on the unfixed spec.

Closes #390

* fix(scrape): correct the 14-day scoping cross-reference, restore EOF newline

Review follow-up on #391.

The 14-day scoping is Step 1b's list item 3, not Step 3 - Step 3 is Quick Fit
Assessment and never touches dates. The "3." list item had been promoted to a
step number. Corrected in the new SKILL.md paragraph (both occurrences), the
CHANGELOG entry, and the test class docstring; a wrong pointer in a file agents
execute as instructions actively misleads.

Also restores the trailing newline on tests/test_scrape_contract.py (the nit
left for a future touch in #344) and adds the (#390) ref to the CHANGELOG entry
to match its siblings.
2026-08-30 20:26:37 +02:00
sdrarunvarshan dea8140db2 feat(ats): extract PDF text with pypdf before Poppler (#369)
* feat(ats): extract PDF text with pypdf before Poppler

Lead the ATS text-layer check with pypdf (BSD, optional pip install). Fall back to pdftotext -layout -enc UTF-8. No cache directory, no installer, no AGPL pymupdf. Windows users without Poppler still get a mechanical parseability check; visual review remains the last resort.

* Update verify_pdf.py

* Update apply.md

* Update verify_pdf.py

* Update verify_pdf.py
2026-08-26 20:07:03 +02:00
Prince chukwuemekaandCursor d82df2fe51 fix(reset): clear the two personalized skill files /reset profile missed (#364) (#365)
/setup Step 3 populates six skill files; /reset profile cleared four.
04-job-evaluation.md was listed by name under "files NOT touched (they
contain framework rules, not candidate data)" while Step 3.4 writes the
user's match areas, career goals, energizing/draining tasks, financial
situation and schedule constraints into it - and ci.yml's
placeholder-integrity job already guards that file under "personal data
may have been committed". job-scraper/search-queries.md, which Step 3.8
fills with their job boards, role titles, domain keywords, city and
commute tiers, appeared nowhere in reset.md at all.

Both are tracked and unignored, so Step 1 asked the user to confirm a
wipe list that omitted them and Step 4 then reported a blank profile
while /rank kept scoring against the old skills and career goals and
/scrape kept running the old city and queries. Re-running /setup does
not necessarily clean them either: Path A skips files whose content is
"no longer placeholder text", and Step 3.8 is phrased as token
replacement, with no tokens left to replace.

Both files are now previewed and cleared, restoring their /setup
placeholders while preserving the scoring framework and the query
structure. 04-job-evaluation.md leaves the preserved list, which keeps
03-writing-style.md and 06-cover-letter-templates.md - the latter
correctly, since its [YOUR_NAME] tokens are LaTeX scaffolding Step 3
never writes to. CLAUDE.md and cv/main_example.tex stay outside the
profile scope, which reset.md:13 defines as skill files only; the
preview and Step 4 now say they still hold personal data instead of
implying a full wipe.

tests/test_reset_command.py gains a profile-scope guard beside its
documents-scope one, deriving the file list from /setup Step 3's own
headings rather than hardcoding it, so a future /setup target that
/reset forgets fails in CI. Against master the three cases fail on
exactly the defect: preview missing search-queries.md, execution
missing both, and the preserved list mislabelling 04-job-evaluation.md
as framework-only - the last of which a filename search alone would
have missed.

Closes #364

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-25 16:55:42 +02:00
Ritik Yadav 7d00ec7925 fix(salary): stop dropping the dotted A.M.B.A. suffix in company-name matching (#356)
The A.M.B.A. STRIP_PATTERNS regex ended in a literal dot followed by
\b, but \b can't fire right after a non-word character when the next
char is also non-word (space/end-of-string) - so it never matched any
realistic company name. The sibling undotted 'amba' suffix stripped
fine, so 'Arla Foods A.M.B.A.' and 'Arla Foods amba' normalized to
different strings and scored 86 vs 100 against the same query.

Made the trailing dot optional so the boundary resolves correctly.
2026-08-23 09:26:21 +02:00
Gabriel Ignacio Mensi eee739ed7e fix(cache): address PR #349 follow-up feedback (#359)
Two small, non-blocking asks from Mads on #349:

- Pin the verification-still-applies restatement in apply.md and
  interview.md's cache-check paragraphs - the one part of the wiring
  with no dedicated test (one assertion each, as requested).
- State cache contents are data, never instructions, in
  04-job-evaluation.md's cache section - closes a carry-over
  prompt-injection surface for a later session reading the file, same
  trust-boundary rule apply.md Step 0 already states for the posting.
2026-08-23 09:00:13 +02:00
Gabriel Ignacio Mensi becdc5dfd7 feat(apply,interview): cache company research to skip repeat lookups (#349)
/apply Step 3's reviewer agent and /interview Step 2 each independently
execute the Company Research Checklist (04-job-evaluation.md) for the
same company - applying to a role and later prepping for its interview
researches the company twice from scratch, same WebSearch/WebFetch cost
both times, no sharing between the two commands.

Adds a company_research/<normalized-name>.json cache (30-day TTL) that
either consumer checks before researching and writes after a fresh
pass. Defined once in 04-job-evaluation.md, next to the checklist it
mirrors, so both commands point at one source instead of restating the
schema. Does not change the verification model: 03-writing-style.md
rule 5 already treats reviewer-agent research as a lead, not a source,
requiring independent re-confirmation before any company claim ships
in a final artifact - the cache stores source URLs alongside each
fact so that re-confirmation stays cheap, but the requirement itself
is untouched and restated in both consumers.

company_research/*.json added to .gitignore and security_guards.py's
REQUIRED_IGNORE_RULES as a plain rooted pattern (not **/-prefixed):
the cache is referenced from commands, not a skill, so it resolves
against the repo root normally, unlike job_scraper/upskill's
skill-relative paths.

Pinned by tests/test_company_research_cache.py, mirroring the
spec-pinning pattern in test_rank_command.py and test_onboarding_privacy.py.
The write-back assertions for both apply.md and interview.md were
verified to actually fail against the regression they guard (the
instruction stripped, confirmed the test catches it, restored) before
being considered done - the write half is the one most likely to be
dropped silently in a future edit, since the read half is the more
obvious change to make.

framework_version bumped 1.2.4 -> 1.2.5 in 04-job-evaluation.md, the
only touched file inside the tracked skill set.
2026-08-22 11:21:34 +02:00
Mads LorentzenandClaude Opus 5 34b8b3f91f fix(onboarding): warn about public forks at the point of decision (#345) (#348)
The quick start walked a new user into gh repo fork - forks of public
repos are always public - and two steps later had /setup write personal
data into tracked files, with the only complete warning in SETUP.md
section 8, a section about pulling updates that a first-time user has no
reason to open during onboarding. A real user hit exactly this (#345).

The warning now sits adjacent to both fork commands (README step 1,
SETUP.md section 2, both pointing at section 8's private-remote recipe),
and /setup checks the origin's visibility BEFORE writing anything: a
public-fork origin gets a confirm-first warning instead of a note after
every file is on disk. A private origin, no origin, or a non-git
directory continues silently. Reported by @basilevs with a complete
reproduction and fix analysis; this implements his fixes (1) and (4).

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 21:33:32 +02:00
Oscar Madera db8312948a test(scraper): pin the /scrape Step 2 search-output contract across portal CLIs (#344)
The Step 2 contract ('Search output already includes title, company,
location, date, and URL') had no cross-portal regression net: a CLI that
quietly drops a contract field flags the portal as degraded on every
/scrape run while CI stays green. That failure class landed for real
(jobnet/jobdanmark/jobbank, fixed in #339/#340/#342). The contract fields
are derived from the SKILL.md sentence itself (never hardcoded), compared
against the real search output of every .agents/skills/*-search CLI, and
detail.ts is deliberately excluded - the contract is about what /scrape
consumes. Fails against pre-fix master on exactly the three portals the
fixes cover; passes with them applied.
2026-08-19 21:23:46 +02:00
Mads LorentzenandClaude Opus 5 9a69309749 fix(scrape): add a client-side recency fallback for flagless portals
Step 1b.3's "scope to 14 days using the portal's recency flag" was
unsatisfiable on jobdanmark, which has no date filter or sort - the
agent either silently skipped the scoping or invented a flag, and the
CLIs now reject invented flags loudly. Every portal emits date, so the
instruction now filters client-side after the call, and stops
presenting --order (a sort) as interchangeable with a filter. Review
finding F32 (2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 20:56:15 +02:00
Mads LorentzenandClaude Opus 5 2c3d2d8558 fix(html-report): funnel from stage history, rejection rate from true rejections
The funnel was computed from current status - a state, not a history -
so an application that interviewed and was later rejected never counted
as reaching Interview, and a hired candidate produced Interview=0. The
stage checkboxes Step 1.2 already merges from outcome.md are the
history; Step 2 and chart 4 now use them. The rejection rate also
counted offer_declined (a success) and withdrawn (candidate-initiated)
as rejections and left unresolved Interview/Offer rows in the
denominator. Review findings F10 and F11 (2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 20:54:46 +02:00
Mads LorentzenandClaude Opus 5 07cec1f227 fix(ci): co-locate placeholder sentinels with the data they guard
cv/main_example.tex's sentinel was [YOUR_NAME] - a header comment and
the pdftitle, neither of which /setup's documented personalization
touches, so CI reported the file clean while it carried a real name,
address, phone and email (the review proved this end to end; the file
is the one CV the gitignore deliberately allows to be committed). The
guard now checks the \name{} and \email{} data lines, and 01's sentinel
moves from the <!-- SETUP comment onto [YOUR_EMAIL]. New test simulates
the /setup edit and requires every checked CV sentinel to be destroyed
by it. Review finding F28 (2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 20:47:51 +02:00
Mads LorentzenandClaude Opus 5 d4e0c64c3c fix(rank): rename the location verdict field to location_verdict
"location" meant a place in scraper output and a PASS/FAIL/FLAG verdict
in /rank's persistence - one key, two meanings, in the same store, with
ranking able to overwrite the commute-filter place with "PASS". The
verdict now lives in location_verdict; legacy entries are read
compatibly and migrated on re-write. Also completes the seen_jobs schema
enumeration (F27 Part A): the do-not-drop instruction now names
location_verdict/language_gate/language_note. Review finding F27
(2026-08-19), decision approved by Mads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 20:36:49 +02:00
Mads LorentzenandClaude Opus 5 65fbe8b8a4 test(framework-version): cover the CI gate that had zero tests
check_framework_version.py guards fork-rebase safety (Gate E) and could
be neutralised by a one-line change that reads as a refactor, with
nothing in the repo noticing - a broken guard is silent by construction.
Four new tests run the real script inside an isolated git repo: clean
tree passes, unbumped edit fails, bumped edit passes, missing marker
fails. Mutation-verified against the exact return-False disable the
review demonstrated. Review finding F22 (2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 20:00:25 +02:00
Mads LorentzenandClaude Opus 5 9a074b262d test(lint-skills): cover check_skill and check_command, not just settings
The linter's main job - frontmatter keys, allowed-tools targets, the
command title rule - had zero assertions; deleting the missing-
allowed-tools error left the suite green. The fixture's yaml stub now
parses the flat frontmatter the fixtures write instead of returning a
canned mapping, and four new cases pin both check functions.
Mutation-verified against the real linter. Review finding F23
(2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:59:22 +02:00
Mads LorentzenandClaude Opus 5 c20458d768 test(robots-check): pin the tie-break clause and the browser-UA fallback
The tie-break test listed Disallow first - the one ordering where
deleting the clause changes nothing - and gate()'s read-the-policy-as-a-
browser recovery (the Barclays-class case 09-web-research.md documents
as covered) had no test. Both gaps are guard code whose breakage is
silent by construction. Mutation-verified: the tie-break deletion and
the UA-loop reduction each now fail exactly the new tests. Review
findings F21 and F30 (2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:57:46 +02:00
Mads LorentzenandClaude Opus 5 4ed5fee221 fix(gmail-sync): replace in:inbox with -in:sent -in:drafts
in:inbox matches only messages currently in the Inbox, so it silently
excluded archived mail and everything routed past the inbox by a
label-and-archive filter - exactly the mail matched by the job-search
label Step 3.1 hunts for. The stated intent ("skip sent/drafts") is what
the negative operators express. Failure mode was silent under-detection
that read as "no updates" and left the tracker stale. Review finding F18
(2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:55:48 +02:00
Mads LorentzenandClaude Opus 5 0e054f16e7 fix(upskill): give Step 3.3 a rule for blank fit_rating rows
/outcome-created tracker rows (applications made outside the workflow)
never got a fit evaluation, so fit_rating is blank - and Step 3.3's
weight formula divides by it with no stated rule. Blank read as 0 means
weight 1.0, the maximum: the job the framework knows least about would
dominate the heatmap and the learning plan. Blank now falls back to a
matched ranked entry's rank_score, else skip+count+report once - the
same pattern the skill already applies to missing gaps. Review finding
F29 (2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:54:32 +02:00
Mads LorentzenandClaude Opus 5 57e82d2b59 fix(rank): treat non-ISO stored deadlines as absent in urgency and the sweep
Rule 6's expiry sweep mutates status automatically from stored deadline
values, yet had no rule for the non-ISO shapes portals have shipped into
seen_jobs.json ("ASAP", DD.MM.YYYY, free text) - "ASAP" is incomparable
and "01.09.2026" is ambiguous between 1 Sep and 9 Jan. /outcome, which
merely displays dates, already carried the defensive-parse rule. A
non-YYYY-MM-DD stored value is now handled like an absent one (left
alone, never compared, never guessed at) and reported once with its
portal. Includes the F24-style coupling test. Review finding F17
(2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:53:29 +02:00
Mads LorentzenandClaude Opus 5 9ab697de64 fix(evaluation): update stale Language Gate preamble to reflect tracking
04-job-evaluation.md still said the gate result "is not a field /scrape
or /rank track" - true when the gate was introduced, false since /rank
began persisting language_gate/language_note as shortlist veto fields
and /scrape began surfacing the flag. The authoritative framework file
taught agents the opposite of rank.md's own persistence rule. New
coupling test pins that the section names the tracked fields and never
reverts to the untracked claim. framework_version 1.2.3 -> 1.2.4.
Review finding F24 (2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:52:21 +02:00
Mads LorentzenandClaude Opus 5 2036b97040 fix(reset): include documents/postings/ in the documents scope
/reset's preview, delete block, and scope description all skipped
documents/postings/ - the drop folder for hand-pasted posting text,
documented in documents/README.md and protected as personal data by
security_guards.py - and then asserted "The documents/ folder is now
empty." The new test derives the folder list from the git tree, so any
future drop folder fails it until /reset covers it. Review finding F26
(2026-08-19).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:50:11 +02:00
Mads LorentzenandClaude Opus 5 1c19f6c45f fix(salary-tools): decide number locale by last separator, pair compound headers
Review findings F7 and F8 (2026-08-19):

- parse_numeric_cell's both-separators branch always assumed European
  locale, silently turning a US "1,234.56" into 1.23456 - a 1000x
  corruption written to salary_data.json with no warning. The separator
  that appears last is now treated as the decimal separator, which also
  makes multi-group values ("1,234,567.89") parse instead of raising a
  raw float error. Single-separator ambiguity guards are unchanged.

- strip_type_patterns stripped only whole tokens, so the compound header
  "Lønindeks alle" kept its type word and could never pair with "Antal
  alle" - failing exactly for the compound-word locale COMPOUND_PATTERNS
  exists to support. It now also strips compound patterns as substrings,
  mirroring header_matches.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:48:50 +02:00
Mads LorentzenandClaude Opus 5 b2545d5121 fix(latex): brace bracket-leading bullets, document escapes, pin pdftotext encoding
Three findings from the 2026-08-19 review (F9, F31, F34):

- F9: every placeholder bullet written as \item [text] let LaTeX parse
  the bracketed text as the item's optional label, rendering it clipped
  off the left page edge and absent from the PDF text layer ("Achievement"
  appeared 9 times in cv/main_example.tex and 0 times in the extraction,
  with a clean compile and green CI). Bullets are now braced as
  \item {[text]} in the example CV and in the template
  06-cover-letter-templates.md teaches, and CI's stock PDF assertions
  additionally require "Achievement" to survive pdftotext.

- F31: 05-cv-templates.md gains a "LaTeX Special Characters" section and
  06's is completed beyond \_ and \&. The load-bearing case is an
  unescaped % in a quantified achievement bullet: it starts a LaTeX
  comment and silently deletes the rest of the line from the PDF.

- F34: the documented ATS extraction commands (apply.md,
  05-cv-templates.md, CLAUDE.md) now carry -enc UTF-8. Xpdf-based
  pdftotext builds default to Latin-1 output, so a correct non-ASCII CV
  failed the replacement-character parseability check.

framework_version: 05-cv-templates.md 1.4.1 -> 1.4.2,
06-cover-letter-templates.md 1.0.1 -> 1.0.2. All three pinned by the new
tests/test_latex_guidance.py (9 tests; suite now 261).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:44:36 +02:00
Yash Rajeshbhai Darji f136b534de feat(scraper): record whether each seen job came from a CLI or the WebSearch fallback (#338)
Adds an additive source field (cli/websearch) to the seen_jobs Step 4 schema, Step 1c tagging at collection time, and a 'fallback (websearch):' Step 5 summary line - so future ghost-job reports (#331) self-triage from stored state. Same additive-field contract as portal/deadline: never backfilled.

Co-authored-by: yshraj <87583119+yshraj@users.noreply.github.com>
2026-08-18 21:06:16 +02:00
Ayobami Adegokeandayobamiseun 2ff1085254 fix(archive): derive <company>_<role> as a single path component (#329)
Extends the canonical Subfolder-naming rule by citation to all six archive derivation sites (apply, gmail-sync, interview, notion-sync, outcome, assistant SKILL.md), adds a fail-closed guard for an empty derived name, and pins every site with mutation-verified tests. framework_version 1.3.3 -> 1.3.4.

jakob1379 independently specified the same fix in his fork's issue #22 before this PR's rework.

Co-authored-by: ayobamiseun <66267222+ayobamiseun@users.noreply.github.com>
2026-08-18 21:03:16 +02:00
Oscar MaderaandJakob Stender Guldberg 762d3218ef fix(html-report): read and render the tracker deadline column (#325)
/html-report was the one tracker consumer #319's deadline column left behind:
Step 1 now parses every canonical column and Step 3 renders Deadline after Date.
The drift guard derives CANONICAL_HEADER from apply.md itself, so a future
column added elsewhere but missing here fails with the column named; legacy
13-field rows read as empty deadline, never dropped, never inferred. Includes
rule-6 sweep refinements and deadline-reconciliation rules authored by
jakob1379.

Co-authored-by: Jakob Stender Guldberg <17257805+jakob1379@users.noreply.github.com>
2026-08-16 20:02:09 +02:00
Oscar MaderaandJakob Stender Guldberg c855e11d22 fix(tracker): persist the application deadline end to end (#319) (#324)
The deadline is written at every moment it is provably in hand and survives
every write that follows: seen_jobs.json base field, /rank stored-value urgency
+ expiry sweep with Step 4 persistence, tracker 14th column with header-line-only
migration for existing files, /scrape-path extraction (assistant SKILL.md 1.3.2
-> 1.3.3), preserve-unparsed-fields in /outcome and /gmail-sync, notion-sync
deadline precedence. Design, scope analysis, and the folded refinements by
jakob1379 (#319, #328).

Co-authored-by: Jakob Stender Guldberg <17257805+jakob1379@users.noreply.github.com>
2026-08-16 19:56:45 +02:00
Oscar Madera 621ce5ab39 fix(convert-salary-excel): reject ambiguous dot thousands separators (#326) 2026-08-14 10:51:06 +02:00
670d30ae7e feat(upstream): add commit-level triage and weekly watch for forks (#305) (#320)
Forks tracking this template face a weekly "which of these commits do I
actually care about?" question. check_upstream_updates.py answers it at the
file level (version stamps); this adds the commit-level half.

tools/upstream_triage.py walks the commits a fork is behind and splits them
into "worth reviewing" and "probably skip". Work already ported drops off on
its own via git patch-id, commits touching only files the fork removed are set
aside, and SHAs in .github/upstream-wontport.txt stay hidden. It reports and
nothing more - ready-to-run cherry-pick lines, but no merge, push, or PR,
since on a fork "applies cleanly" is not "correct".

.github/workflows/upstream-watch.yml runs it weekly into one rolling issue. It
no-ops on the upstream template (guarded, and pinned by a test) and uses only
the built-in GITHUB_TOKEN, so it can never write outside its own fork. The two
tools point at each other in their output; README, SETUP 8, and CHANGELOG
introduce them together. Tests cover patch-id matching, relevance filtering,
the won't-port list, and the workflow guard - all offline.

Co-authored-by: Angelina Lok <angelina@chattermill.io>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-11 19:11:24 +02:00
Ayobami Adegoke cfd9a9fba1 fix(security): gitignore /upskill reports at the path the skill writes them (#317)
The upskill/*.md rule is rooted, but /upskill is a skill and skills
resolve bare relative paths against their own directory - the same
observed behavior the **/job_scraper rules exist for. A report at
.claude/skills/upskill/upskill/report-*.md was not ignored, and an
upskill report records the candidate's skill gaps against named
employers. **/upskill/*.md would also ignore the skill's own SKILL.md
(the directory shares the name), so the new rule pins the report-file
prefix: **/upskill/report-*.md. Added to .gitignore and
REQUIRED_IGNORE_RULES, with a check-ignore-based test pinning both
properties.
2026-08-11 19:07:21 +02:00
Muhammad Haseeb 3efc52ebd5 feat(security): hold .claude/settings.json hooks to a reviewed allowlist (#313)
check_permissions() read permissions.allow and nothing else, so a `hooks`
block in the same file passed the guard silently.

A hook is strictly more dangerous than a pre-approved permission. A
permission pre-approves something Claude may choose to do; a hook runs
unconditionally when its event fires, with no prompt and no model decision
in between. Cloning the repo and opening it is enough.

This is the vector the Shai-Hulud worm used in its August 2026 wave: a
SessionStart hook in .claude/settings.json chaining to .claude/math_init.js,
executing on session start.
https://research.jfrog.com/post/shai-hulud-is-back-august/

For a template that thousands of people are explicitly invited to fork,
that is the riskiest key in the file this guard already parses.

Follows the established pattern exactly - ALLOWED_HOOKS ships empty, since
the template has no hooks, so any addition must be allowlisted in the same
PR and is therefore explicit and reviewable.

Two details worth reviewing closely:

- The hook check runs *before* the permissions shape guards. Those guards
  return early, so a file pairing a malformed permissions block with a live
  hook would otherwise skip the hook check entirely - a fail-open. Pinned by
  test_hook_is_caught_even_when_permissions_block_is_malformed.
- _hook_commands() fails closed. Any hook layout it does not recognise
  yields a marker that cannot be in the allowlist, so an unfamiliar shape is
  rejected rather than silently skipped, rather than trusting that the
  Claude Code schema will not change.

Verified:
  - 8 new HookGuardTests cases; 14 of the suite's 26 tests fail against the
    unpatched guard, all 26 pass with it
  - injecting the real worm shape into this repo's own settings.json makes
    the guard exit 1 naming 'SessionStart:node .claude/math_init.js';
    removing it returns OK
  - lint_skills, check_framework_version, security_guards all OK;
    python3 -m unittest discover -s tests 219 passed
2026-08-10 21:18:33 +02:00
Jakob Stender Guldberg 0e1a895c4e fix(apply): archive the job posting while /apply still holds it (#306) (#307)
/apply drafted two documents and a tracker row from the full posting, then
let the text die with the session. /outcome Step 3.2 tried to recover it by
re-fetching a `source` URL the spec itself expects to be dead, and a posting
pasted from an email or a PDF had no `source` to re-fetch at all.

Step 6b gains item 7: write the posting verbatim to
documents/applications/<company>_<role>/job_posting.md, never a re-fetch or a
reconstruction from memory. The folder is derived by citing /outcome Step 1.4
rather than restating the rule, so the two cannot drift. An existing file is
left alone and named in the report.

Step 0 and the /scrape path (job-application-assistant SKILL.md Step 1) now
retain the full posting text rather than a summary, so item 7 has something
verbatim to write.

Pinned by tests/test_apply_records_application.py.
2026-08-09 20:26:58 +02:00