An ATS reads the compiled PDF's embedded text layer, not the rendered page,
and LaTeX can silently produce PDFs whose text extracts as garbage: icon
glyphs where contact details should be, (cid:*) markers from fonts without
Unicode mappings, interleaved lines from multi-column layouts. This matters
more now that /add-template lets users bring arbitrary templates. The
existing Step 5 loop verifies what a human sees; this adds verification of
what a parser sees.
New Step 5d in /apply (CV only - cover letters rarely go through keyword
screening; cleanup renumbered to 5e):
- Extract the CV PDF's text layer with pdftotext -layout. pdftotext
(poppler) is an optional dependency: if missing, the mechanical check is
skipped with a warning and keyword coverage falls back to the visual PDF
read - the same graceful-skip pattern as salary_lookup.py
- Parseability checks verified against a real extraction of the stock
template: email/phone must survive as literal text (fontawesome icons
extract as harmless glyph-name noise like MOBILE-ALT/Envelope, but a
contact detail carried only by an icon or hyperlink is invisible to ATS),
no (cid:*) or replacement-character garbage, reading order matching
visual order, dates present
- Keyword coverage reuses the required/preferred list from Step 1, matched
in the posting's language, reported as covered / synonym-only /
missing-have-it / missing-gap. Honesty rule enforced: keywords the
profile genuinely supports get added to experience bullets; genuine gaps
stay visible, never stuffed
Integration: CLAUDE.md verification checklist section, ATS Parseability
guidance in 05-cv-templates.md, narrow Bash(pdftotext:*) entry in the
pre-approved permissions (keeping with the tightened scope from #27),
cv/*.txt gitignored (extraction is personal data; also deleted by the
step itself), and optional-dependency docs in README and SETUP.
The trailing `.agents/` rule ignored the entire job-search CLI tree, so the
scraper skills (the existing Danish ones, and any portal skill a forker adds)
are silently excluded from the repo unless force-added. Narrow it to ignore
only node_modules, logs, and usage data, so skill source is tracked.
Co-authored-by: Akhil Tripathi <kodabear@Akhils-MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Adds four contributions from @Michael-Bach:
- /setup_docs - document-driven profile population from a documents/ folder (CV, LinkedIn export, diplomas, references, past applications). Idempotent merge with explicit additive vs. conflicting buckets and per-conflict prompts.
- /reset - typed-RESET confirmation gate for clearing profile data and/or documents folder contents.
- /expand - additive competency enrichment from documents and public URLs already in the profile (GitHub repos, portfolio sites), with web-searched syllabus lookups for named courses and certifications.
- /upskill - skill-gap analysis vs tracked jobs (or single URL), produces a prioritized heatmap and learning plan with year-tagged WebSearch queries.
Also adds the documents/ folder convention with README and gitignore entries for personal output files.
Closes#6