diff --git a/.claude/skills/job-application-assistant/05-cv-templates.md b/.claude/skills/job-application-assistant/05-cv-templates.md index 749d49e..f25c457 100644 --- a/.claude/skills/job-application-assistant/05-cv-templates.md +++ b/.claude/skills/job-application-assistant/05-cv-templates.md @@ -1,5 +1,5 @@ --- -framework_version: 1.4.3 +framework_version: 1.4.4 --- # CV Templates and Tailoring Guide @@ -42,6 +42,13 @@ Expected output: `Output written on main__.pdf (2 pages, ...)`. A \renewcommand*{\sectionstyle}[1]{{\sectionfont\color{color1}#1}} \usepackage[utf8]{inputenc} +% pdflatex fallback only (the documented engine is lualatex, which skips this +% branch). Without T1 font encoding pdflatex builds accented letters with +% \accent, and the PDF text layer stores them decomposed - `e` + U+0300 rather +% than U+00E8 - so an ATS keyword match on "Genève" fails while the page looks +% right. moderncv 2.5 loads T1 itself under pdflatex; 2.3.1 (Debian/Ubuntu apt) +% does not. \ifpdftex comes from iftex, which every moderncv version loads. +\ifpdftex\usepackage[T1]{fontenc}\fi % moderncv loads hyperref itself in an \AtEndPreamble hook, so \hypersetup % must go in an \AtEndPreamble of our own: on moderncv < 2.4 a top-level % \usepackage{hyperref} clashes with the class's own @@ -277,7 +284,8 @@ What to check in the extraction: - **Contact details as literal text.** The stock template's fontawesome contact icons extract as glyph names (`MOBILE-ALT`, `Envelope`) - harmless noise, because the actual address and number are printed beside them. The failure mode is a contact detail carried *only* by an icon or a hyperlink (like the `LinkedIn` link text, whose URL is not in the text layer): invisible to an ATS. The email address must always appear as printed text. - **No garbled output.** `(cid:NNN)` markers or `�` characters mean a font is embedded without a Unicode mapping - an ATS sees the same garbage. This shows up with unusual fonts in custom templates, not with the stock moderncv setup under lualatex. - **Reading order.** The stock banking style is single-column, so extraction order matches visual order. Custom templates (via `/add-template`) with sidebars or multi-column layouts can interleave unrelated lines; if extraction order is scrambled, the user is trading ATS compatibility for looks and should be told. -- **Keyword coverage.** Match the posting's required/preferred terms against the extracted text, in the posting's language. Prefer the posting's exact term over a synonym when it is truthfully applicable - ATS matching is often literal. Never add a keyword the profile does not support. +- **Keyword coverage.** Match the posting's required/preferred terms against the extracted text, in the posting's language. Prefer the posting's exact term over a synonym when it is truthfully applicable - ATS matching is often literal. Never add a keyword the profile does not support. `verify_pdf.py --contains` folds both sides for whitespace, Unicode normalization (NFC) and LaTeX's typographic substitutions before comparing - `'` reaches the text layer as U+2019 and `--` as U+2013, so `--contains "Master's degree"` and `--contains "2016-2024"` match what the template actually renders. The dumped `.txt` is never folded: it is the raw layer the ATS sees, which is why the date-range check below reads the dump, not `--contains`. +- **Accents intact (pdflatex fallback).** Under pdflatex without T1 font encoding the text layer stores accented letters decomposed (`e` + combining grave instead of `è`); pypdf reads that as `Gen` `eve` with a stray spacing accent, and neither form matches a typed keyword. The stock template guards this with `\ifpdftex\usepackage[T1]{fontenc}\fi`; keep the line in tailored CVs and custom templates that may be compiled with pdflatex. It is a no-op under lualatex. ### Date fields must be ASCII ranges (confirmed ATS import failure) diff --git a/CHANGELOG.md b/CHANGELOG.md index 8a967dc..ad3b830 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -47,6 +47,26 @@ per-file diff commands. ### Fixed +- **`verify_pdf.py --contains` now sees through LaTeX's typographic substitutions and + the pdflatex text layer keeps accents precomposed** (Discussions #385, #384) - the + comparison folded whitespace only, but LaTeX ligatures `'` into U+2019 and `--` into + U+2013, so on the stock CV compiled with the documented `lualatex` command + `--contains "Master's degree"` and `--contains "2016-2024"` both reported the keyword + missing from a document that plainly contains it (measured through both extractors; + `Six Sigma` and `Statistics` on the same page passed). The documented remedy for a + missing keyword is to add it, so the false negative nudged toward the one thing the ATS + section forbids. `normalize_text()` now folds both sides - NFC, then curly + apostrophes/quotes to ASCII, en/em dashes to `-`, no-break space to space - at + comparison time only; `--dump-text` still writes the raw layer, because that is what an + ATS parses and the date-range rule in `05-cv-templates.md` needs the raw en-dash visible + there. Separately, pdflatex without T1 font encoding stores accents decomposed + (`e` + U+0300; pypdf reads it as a stray spacing accent), which NFC cannot fully + repair - moderncv 2.5 loads T1 itself under pdflatex but the apt-packaged 2.3.1 does + not, so `cv/main_example.tex` and the guide's preamble gain + `\ifpdftex\usepackage[T1]{fontenc}\fi`, a no-op on the lualatex path. Pinned by + ten new `test_verify_pdf.py` cases (the fold-through ones fail on the whitespace-only + code) and a `test_latex_guidance.py` guard that the line exists and stays + pdflatex-only. Reported and diagnosed by 9scorp4. - **`jobdanmark-search detail` now backs off on 429/5xx like every other portal's detail command** - the handler called `fetch()` directly instead of going through the CLI's own request wrappers, so it carried none of the three things `apiFetch`/`apiPost` guarantee: diff --git a/cv/main_example.tex b/cv/main_example.tex index 59bb7f3..f6d786d 100644 --- a/cv/main_example.tex +++ b/cv/main_example.tex @@ -23,6 +23,13 @@ \renewcommand*{\sectionstyle}[1]{{\sectionfont\color{color1}#1}} \usepackage[utf8]{inputenc} +% pdflatex fallback only (the documented engine is lualatex, which skips this +% branch). Without T1 font encoding pdflatex builds accented letters with +% \accent, and the PDF text layer stores them decomposed - `e` + U+0300 rather +% than U+00E8 - so an ATS keyword match on "Genève" fails while the page looks +% right. moderncv 2.5 loads T1 itself under pdflatex; 2.3.1 (Debian/Ubuntu apt) +% does not. \ifpdftex comes from iftex, which every moderncv version loads. +\ifpdftex\usepackage[T1]{fontenc}\fi % moderncv loads hyperref itself in an \AtEndPreamble hook, so \hypersetup % must go in an \AtEndPreamble of our own: on moderncv < 2.4 a top-level % \usepackage{hyperref} clashes with the class's own diff --git a/tests/test_latex_guidance.py b/tests/test_latex_guidance.py index ca8539a..c5e85d0 100644 --- a/tests/test_latex_guidance.py +++ b/tests/test_latex_guidance.py @@ -136,5 +136,45 @@ class TestAtsExtractionEncoding(unittest.TestCase): self.assert_pdftotext_commands_pin_utf8(CV_TEMPLATES) +class TestPdflatexFontEncodingGuard(unittest.TestCase): + """#384: the pdflatex fallback must load T1 fontenc, and only under pdflatex. + + Without T1, pdflatex stores accented letters decomposed in the text layer + (`e` + U+0300), so an ATS keyword match on `Genève` fails while the PDF + looks right. moderncv 2.5 loads T1 itself; the apt-packaged 2.3.1 does not. + The line must be guarded so the documented lualatex path is untouched. + """ + + GUARDED_FONTENC = re.compile(r"\\ifpdftex\s*\\usepackage\[T1\]\{fontenc\}\s*\\fi") + + def assert_has_guarded_fontenc(self, path): + text = path.read_text(encoding="utf-8") + self.assertRegex( + text, + self.GUARDED_FONTENC, + f"{path.name} must carry `\\ifpdftex\\usepackage[T1]{{fontenc}}\\fi` so a " + "pdflatex fallback keeps accents precomposed in the text layer", + ) + unguarded = [ + f"{path.name}:{lineno}: {line.strip()}" + for lineno, line in enumerate(text.splitlines(), 1) + if "fontenc" in line + and not line.lstrip().startswith("%") + and not self.GUARDED_FONTENC.search(line) + ] + self.assertEqual( + unguarded, + [], + "fontenc must stay inside the \\ifpdftex guard - lualatex output " + "must not change:\n" + "\n".join(unguarded), + ) + + def test_example_cv_guards_fontenc_for_pdflatex(self): + self.assert_has_guarded_fontenc(EXAMPLE_CV) + + def test_cv_guide_preamble_guards_fontenc_for_pdflatex(self): + self.assert_has_guarded_fontenc(CV_TEMPLATES) + + if __name__ == "__main__": unittest.main() diff --git a/tests/test_verify_pdf.py b/tests/test_verify_pdf.py index 46e772b..6cb5bd2 100644 --- a/tests/test_verify_pdf.py +++ b/tests/test_verify_pdf.py @@ -7,6 +7,7 @@ from unittest.mock import patch from tools.verify_pdf import ( VerificationError, extract_text_layer, + normalize_text, parse_page_count, run_tool, verify_pdf, @@ -22,6 +23,44 @@ class ParsePageCountTests(unittest.TestCase): parse_page_count("Title: Example\n") +class NormalizeTextTests(unittest.TestCase): + """`--contains` must see through what LaTeX does to plain source text. + + Measured on the stock CV compiled with the documented `lualatex` command + (#385): the apostrophe in `Master's` reaches the text layer as U+2019 and + the `--` in `2016--2024` as U+2013, so a whitespace-only fold reports both + keywords missing from a CV that plainly contains them. Under pdflatex + without T1 font encoding, accents arrive decomposed (`e` + U+0300) instead + of precomposed (#384). + """ + + def test_folds_curly_apostrophe_to_ascii(self): + self.assertEqual(normalize_text("Master\u2019s degree"), "Master's degree") + self.assertEqual(normalize_text("\u2018quoted\u2019"), "'quoted'") + + def test_folds_curly_double_quotes_to_ascii(self): + self.assertEqual(normalize_text("\u201cSix Sigma\u201d"), '"Six Sigma"') + + def test_folds_en_and_em_dashes_to_hyphen(self): + self.assertEqual(normalize_text("2016\u20132024"), "2016-2024") + self.assertEqual(normalize_text("role\u2014title"), "role-title") + + def test_folds_no_break_space_to_space(self): + self.assertEqual(normalize_text("EUR\u00a0600k"), "EUR 600k") + + def test_folds_decomposed_accents_to_nfc(self): + decomposed = "Gene\u0300ve Universite\u0301" + precomposed = "Gen\u00e8ve Universit\u00e9" + self.assertEqual(normalize_text(decomposed), precomposed) + + def test_still_collapses_whitespace(self): + self.assertEqual(normalize_text("Professional\n Experience "), "Professional Experience") + + def test_fold_is_symmetric(self): + # A user who pastes the curly form from a posting must match an ASCII layer too. + self.assertEqual(normalize_text("Master\u2019s"), normalize_text("Master's")) + + class VerifyPdfTests(unittest.TestCase): def setUp(self): self.temp_dir = tempfile.TemporaryDirectory() @@ -73,6 +112,41 @@ class VerifyPdfTests(unittest.TestCase): with self.assertRaisesRegex(VerificationError, "Professional Experience"): verify_pdf(self.pdf, required_text=("Professional Experience",)) + @patch("tools.verify_pdf._extract_pypdf", return_value=None) + @patch("tools.verify_pdf.run_tool") + def test_required_text_matches_latex_typographic_substitutions(self, mock_run_tool, _pypdf): + # What the stock template's lualatex text layer actually contains for the + # source `Master's degree ... 2016--2024` (code points measured, see class + # docstring of NormalizeTextTests). + mock_run_tool.side_effect = [ + "Master\u2019s degree in Statistics. Six Sigma Green Belt, 2016\u20132024.\n", + "Pages: 1\n", + ] + + verify_pdf(self.pdf, required_text=("Master's degree", "2016-2024")) + + @patch("tools.verify_pdf._extract_pypdf", return_value=None) + @patch("tools.verify_pdf.run_tool") + def test_required_text_matches_decomposed_pdflatex_accents(self, mock_run_tool, _pypdf): + mock_run_tool.side_effect = [ + "Universite\u0301 de Gene\u0300ve\n", # pdflatex without T1 fontenc + "Pages: 1\n", + ] + + verify_pdf(self.pdf, required_text=("Universit\u00e9 de Gen\u00e8ve",)) + + @patch("tools.verify_pdf._extract_pypdf", return_value=None) + @patch("tools.verify_pdf.run_tool") + def test_dump_text_keeps_the_raw_layer_unfolded(self, mock_run_tool, _pypdf): + # The fold is comparison-time only: the ATS parser sees the raw layer, and + # the date-range rule in 05-cv-templates.md needs the en-dash visible here. + mock_run_tool.side_effect = ["2016\u20132024\n", "Pages: 1\n"] + dump = Path(self.temp_dir.name) / "dump.txt" + + verify_pdf(self.pdf, required_text=("2016-2024",), dump_text=dump) + + self.assertEqual(dump.read_text(encoding="utf-8"), "2016\u20132024\n") + def test_rejects_missing_pdf(self): with self.assertRaisesRegex(VerificationError, "PDF does not exist"): verify_pdf(Path(self.temp_dir.name) / "missing.pdf") diff --git a/tools/verify_pdf.py b/tools/verify_pdf.py index 5349b8e..25fa2b3 100644 --- a/tools/verify_pdf.py +++ b/tools/verify_pdf.py @@ -4,12 +4,18 @@ Text-layer extraction tries pypdf (BSD, optional `pip install pypdf`) first, then Poppler `pdftotext` if pypdf is missing, raises, or returns zero extractable characters. Poppler remains the fallback. + +`--contains` compares after `normalize_text()` has folded both sides: whitespace, +Unicode normalization form (NFC), and the typographic substitutions LaTeX makes to +the source text. The fold is comparison-time only - the `--dump-text` output stays +the raw text layer an ATS parser actually sees. """ import argparse import re import subprocess import sys +import unicodedata from pathlib import Path @@ -47,7 +53,34 @@ def parse_page_count(pdfinfo_output): return int(match.group(1)) +# Typographic substitutions the moderncv/cover.cls templates produce from plain +# source text, mapped back to what a user types into --contains. LaTeX ligatures +# ' into U+2019 and -- into U+2013, so "Master's degree" and "2016-2024" are +# absent from the text layer of a CV that plainly contains them (#385). Applied +# to both sides of the comparison; the extracted dump is never rewritten. +TYPOGRAPHIC_FOLDS = str.maketrans( + { + "\u2018": "'", # ` -> quoteleft + "\u2019": "'", # ' -> quoteright (the possessive apostrophe) + "\u201c": '"', # `` -> quotedblleft + "\u201d": '"', # '' -> quotedblright + "\u2013": "-", # -- -> endash (the \cventry date-range case) + "\u2014": "-", # --- -> emdash + "\u00a0": " ", # ~ -> no-break space + } +) + + def normalize_text(text): + """Fold a string for comparison: NFC, typographic punctuation, whitespace. + + NFC covers the pdflatex text layer, which without T1 font encoding stores + accented letters decomposed (`e` + U+0300) while a user types them + precomposed (U+00E8); both forms fold to the same string (#384). The fold + applies to what is compared, never to what is dumped: the date-range rule in + `05-cv-templates.md` still needs the raw en-dash visible in `--dump-text`. + """ + text = unicodedata.normalize("NFC", text).translate(TYPOGRAPHIC_FOLDS) return " ".join(text.split()) @@ -144,7 +177,11 @@ def build_parser(): "--contains", action="append", default=[], - help="text that must appear after whitespace normalization; repeatable", + help=( + "text that must appear in the text layer; both sides are folded for " + "whitespace, NFC, and LaTeX's typographic substitutions (curly " + "apostrophes/quotes, en/em dashes, no-break spaces); repeatable" + ), ) parser.add_argument( "--dump-text",