3.8 KiB
Salary Benchmark Tool
What is this?
The salary lookup tool (salary_lookup.py) lets you benchmark company salaries against a baseline from your own data. It's used during the /apply workflow to show how a company's compensation compares to market rates.
This tool is optional. If you don't have salary data, the salary step is simply skipped during /apply.
How it works
The tool reads a salary_data.json file in the repo root containing company salary benchmarks. It uses fuzzy matching to find companies by name, handling Danish/Nordic characters, legal suffixes (A/S, ApS), and common spelling variations.
The data format supports any index-based or absolute salary data. For example:
- Index 100 = median salary, higher is better
- Absolute salary values in your currency
- Any custom metric you want to track
Data format
The tool expects salary_data.json with this structure:
{
"metadata": {
"source": "My Union Statistics 2025",
"index_baseline": 100,
"index_label": "Index",
"baseline_description": "Index 100 = median salary for private sector"
},
"companies": [
{
"company": "Novo Nordisk A/S",
"city": "Bagsværd",
"categories": {
"all_employees": { "count": 500, "index": 108.5 },
"engineering": { "count": 120, "index": 112.3 }
}
},
{
"company": "Ørsted A/S",
"city": "Fredericia",
"categories": {
"all_employees": { "count": 200, "index": 105.2 }
}
}
]
}
Fields
- metadata.source: Where the data comes from (for reference)
- metadata.index_baseline: The baseline value (e.g., 100 for index-based data)
- metadata.index_label: Label for the index column in output
- metadata.baseline_description: Human-readable explanation of the baseline
- companies[].company: Company name (required)
- companies[].city: City/location (optional, used for filtering)
- companies[].categories: Named salary categories, each with
countand/orindex
Setup options
Option A: Create salary_data.json manually
Create the file by hand with data from any source: union statistics, Glassdoor, salary surveys, networking, or personal research.
Option B: Convert from Excel
If you have salary data in an Excel file:
pip install openpyxl
python tools/convert_salary_excel.py path/to/salary-data.xlsx \
--source "My Salary Data 2025" \
--baseline 100 \
--baseline-desc "Index 100 = median salary"
The converter auto-detects the Excel layout:
- Looks for a "Company"/"Firma" column and an optional "City"/"By" column
- Treats remaining columns as salary data (auto-pairs count/index columns)
Option C: Build from research
Start with an empty template and add companies as you research them:
{
"metadata": {
"source": "Personal research",
"index_baseline": 0,
"index_label": "Monthly salary (DKK)",
"baseline_description": "Approximate monthly salary before tax"
},
"companies": [
{
"company": "Example Corp",
"city": "Copenhagen",
"categories": {
"entry_level": { "index": 42000 },
"senior": { "index": 55000 }
}
}
]
}
Usage
python salary_lookup.py "Novo Nordisk"
python salary_lookup.py "Ørsted" --city "Fredericia"
python salary_lookup.py "COWI" --json
python salary_lookup.py --list-all
Important notes
- The data file (
salary_data.json) is excluded from git (see.gitignore). Your salary data may be proprietary or confidential. - If the data file is missing,
salary_lookup.pyexits with a helpful error message and the/applyworkflow skips the salary benchmark step. - The fuzzy matcher handles Danish company name variations: legal suffixes, Nordic characters, anglicized spellings, and partial matches.