# How Helix works out salary and working-life information

This document explains where every salary figure in Helix comes from, how firmly
each one is grounded, and what it does and does not mean. It is written to be
read by anyone assessing whether the numbers can be trusted — no programming
knowledge is needed.

The short version: Helix publishes a salary range for all 716 careers, but it
does not pretend they are all equally well evidenced. **148 come from an official
careers guide written for that exact job. The other 568 are estimates, and every
one of them is labelled as an estimate wherever it appears.**

---

## 1. What the numbers are, and what they are not

A Helix salary range is **a decision-support estimate for comparing careers**. It
is a description of what a career typically pays in the UK.

It is **not**:

- an offer, or a prediction of what you personally would be paid
- a statement about a particular employer, region or grade
- a guarantee that any job at that salary currently exists

Pay varies substantially by employer, sector, location, experience, hours and
working pattern. Two people with the same job title can be paid very differently
and both be normal.

Helix rounds figures to sensible whole amounts. You will see "£30k to £53k", not
"£43,281". False precision is a way of implying research that was never done.

A figure taken directly from an official source is republished exactly as it was
issued. A derived figure is rounded to the nearest thousand, because it is the
output of an average rather than a published number, and its digits should not
claim a precision the arithmetic cannot support.

### What a published range actually spans

A National Careers Service range runs from a **starting salary** to an
**experienced** one, across the whole of a career. It is not the span of a single
pay grade.

That distinction matters most in the NHS. Biomedical Scientist is published as
£30,000 to £53,000, and a biomedical scientist entering the profession is on
Agenda for Change Band 5 — whose top is well below £53,000 nationally. The two
figures are not in conflict: the £53,000 end describes an experienced biomedical
scientist who has progressed, typically into the specialist and senior grades
above Band 5, not somebody at the top of Band 5. Helix now labels the two ends
explicitly on each career page for exactly this reason.

So a higher upper figure does not mean the private sector pays more. It usually
means the range covers more of a career.

Where a role sits on a public-sector pay framework, that framework's own bands
are the right thing to compare a specific post against. Helix is built to show
them alongside the market estimate (never instead of it), but only from a curated
mapping recorded by a person with its source — see the limitations.

---

## 2. The three data files, and why they are separate

| File | What it holds | How often it changes |
|---|---|---|
| `data/careerpath_uk_careers_v1.json` | The 677 supplied UK careers: titles, families, tags, regulation status, official sources | Never. Treated as supplied and verified by hash |
| `data/helix_additional_careers_v1.json` | Careers added since launch, ids from CP-701 | When a career is added |
| `data/helix_market_data_uk_v1.json` | Salary, hours, working patterns and role descriptions, one record per career | Monthly, or whenever a source updates |

**Market data is separate from the career list** because salary is volatile and a
career taxonomy is not. Mixing them would mean rewriting the careers file every
time a pay figure moved, and every rewrite of a hand-built reference file is a
chance to corrupt it. They are joined in memory by career id when the application
loads.

**Added careers are separate from supplied ones** because the supplied file is
treated as immutable: the refresh checks its hash on every run and fails if it
has moved. That is the guarantee that nothing has quietly rewritten the launch
taxonomy, and it is worth more than the tidiness of a single file. Ids from
CP-701 leave a clear gap after the supplied CP-001 to CP-677, so an id says at a
glance where a career came from and an addition can never shadow a supplied
record.

Every tool counts the careers it finds rather than assuming a fixed number, so
the catalogue can grow without anything failing for the wrong reason.

---

## 3. Where the data comes from

All sources are official UK bodies. Helix does not use recruitment marketing
pages, salary blogs, forums or search-engine results as evidence of anything.

### National Careers Service

The National Careers Service publishes a job profile for each of roughly 737
careers, including a salary range for a starter and for an experienced worker,
typical weekly hours, working patterns and a description of the role. This is the
preferred source, because it is written about a specific job rather than a
statistical occupation group.

Helix uses it in two ways: through the Job Profiles API where a subscription key
is configured, and otherwise from the public job-profile pages. Both are Crown
copyright, published under the Open Government Licence v3.0.

**Current status:** the Job Profiles API is subscribed to but not yet reachable —
the developer portal currently publishes no APIs to call, and every candidate
route returns "not found" whether a key is supplied or not. This is a
provisioning matter with the National Careers Service, not a problem with the
key. The pipeline is written against the API and will use it the moment a working
base URL exists; until then it uses the public profiles, which carry the same
published figures.

### Office for National Statistics

ONS Annual Survey of Hours and Earnings data gives earnings by occupation, keyed
to Standard Occupational Classification (SOC) codes. It is the main official
fallback when no career-specific guide exists.

A four-digit SOC code describes a reasonably specific occupation; a two-digit
code describes a broad group. Helix records which was used and grades the
evidence accordingly. A broad group estimate is never presented as though it
described the exact job.

### NHS and public-sector pay frameworks

Where a career has been deliberately mapped to a public-sector pay band by a
person, Helix can show that band as **context alongside** the market estimate,
never instead of it.

Helix will not infer an Agenda for Change band from a job title. Words like
"Senior", "Specialist" and "Manager" mean different things in different
organisations, and guessing a band from one of them would produce an official
looking number with nothing behind it. England, Scotland, Wales and Northern
Ireland are kept separate, because their pay frameworks are separate.

### Skills England

Used for occupational mapping where it links a career to a SOC code. Not used as
a salary source.

---

## 4. How a salary is chosen

Every career goes through the same five steps, in order, and stops at the first
one that produces a defensible answer.

**Step 1 — A careers guide written for this exact job.**
If a National Careers Service profile matches the career by title, and the two
are genuinely compatible, its published range is used directly.
Result: **Career-specific guide**.

**Step 2 — A verified public-sector pay framework.**
Only where somebody has curated the mapping. Used as the main range only when it
is more specific to the role than anything else available.
Result: **Strong estimate**.

**Step 3 — Official occupation earnings.**
ONS earnings for the most detailed SOC code that can be defended.
Result: **Strong estimate** for a specific four-digit mapping, **Indicative
estimate** for a broad group.

**Step 4 — Closely related careers.**
Where no direct source exists, the range is derived from careers that *do* have
stronger evidence and are genuinely similar — measured on shared subject tags,
career family, seniority level and title wording. It uses a similarity-weighted
median across several careers, not the highest or the nearest single one.
Result: **Indicative estimate**.

**Step 5 — Family and seniority median.**
The last resort, and the reason all 716 careers have a figure: a robust median
across careers in the same family at a comparable level.
Result: **Limited-data estimate**.

### How seniority is priced

A derived range is adjusted when the career is a more or less senior grade than
the careers it was derived from. The ladder separates practitioner, specialist,
senior, manager, lead, consultant and executive grades, so that Specialist,
Senior, Lead and Consultant Biomedical Scientist are priced as the four
progressive grades they are rather than all reported at one salary.

One rung is treated with more caution than the rest. §15 warns that "Specialist"
does not mean the same seniority in every sector, and the titles bear it out: a
Specialist Biomedical Scientist is a real grade, an Information Governance
Specialist is a subject-matter role at no particular grade, and nothing in the
title separates them. A "Specialist" title is therefore never allowed to push an
estimate above everything it was derived from.

"Consultant" is read the same way. As a prefix — Consultant Biomedical Scientist,
Consultant in Public Health — it is the senior clinical grade. As a trailing noun
— Quality Consultant, Life Sciences Consultant — it is an advisory role at no
particular grade, and carries no seniority claim.

### Why a good match is sometimes refused

Title matching does not accept a match merely because it would improve coverage.
"Senior Biomedical Scientist" matches the "Biomedical scientist" profile
perfectly on words, but accepting it would publish an entry-grade range as a
career-specific fact about a senior post. Seniority variants are always sent to
derivation instead.

A transparent derived estimate is better than a confident wrong one.

---

## 5. What the evidence labels mean

Every figure in Helix carries one of these, and it is shown next to the number
rather than hidden in a footnote.

| Label | What it means | Count |
|---|---|---|
| **Career-specific guide** | An official careers source published this range for this exact job | 148 |
| **Strong estimate** | A high-quality occupation or pay-framework mapping, but not published for this exact title | 0 |
| **Indicative estimate** | Derived from closely related careers with stronger evidence | 378 |
| **Limited-data estimate** | A median across the career's family and seniority level. A broad indication only | 190 |

There are currently no Strong estimates. That is because the ONS occupation step
requires a defensible SOC mapping, and Helix does not yet have human-verified SOC
codes for these careers. Curating them is the single change that would move the
largest number of careers up a grade, and it is the honest reason the figure is
zero rather than an oversight.

The other route is a curated list of alternative titles for the National Careers
Service profiles, in `data/reference/ncs_career_aliases.json`. 737 public
profiles exist and only 54 matched by exact title, so human-checked aliases
convert derived estimates into career-specific evidence without waiting for
anybody.

**55 have been curated so far**, taking career-specific coverage from 54 to 110.
The effect was larger than those 32 careers: giving derivation better anchors
moved a further 182 careers off the family-median fallback, so Limited-data
estimates fell from 421 to 239.

Each alias is a judgement, made against one test: not "are these titles similar"
but "is the range this profile publishes an honest answer for this career". A
qualifier naming only a sector or setting is usually safe — a marketing manager
in healthcare is a marketing manager. A qualifier marking a different profession,
a different registration or a materially different pay market is not.

**Ten candidates were examined and deliberately rejected**, with the reason
recorded beside each in the alias file so they are not re-litigated. Some
examples: a Clinical Geneticist is a medical consultant, not the laboratory
scientist the *Geneticist* profile describes; a Health Psychologist and a
Clinical Psychologist are different HCPC-protected professions; and *Public
Health Intelligence Analyst* nearly matched the criminal intelligence analyst
profile, which is policing work.

`docs/MARKET-DATA-AUDIT.md` lists whatever candidates remain. Two warnings apply
to that list. A high score is not agreement: matching ignores setting words like
*clinical* and *healthcare*, so *Clinical Photographer* and *Photographer* both
reduce to the same tokens and score 1.00 while plainly being different jobs —
those rows are flagged. And seniority variants are excluded entirely, because
aliasing *Senior Biomedical Scientist* to the entry-grade profile would publish a
starter salary as fact about a senior post.

---

## 6. Working life: hours, patterns and the inferred measures

Two very different kinds of information sit side by side here, and Helix
distinguishes them everywhere they appear.

**Recorded by a source.** Typical weekly hours and working patterns — shifts,
evenings and weekends, on-call, bank holidays — come from official job profiles.
They exist for the 148 careers with a matched profile. For the other 568 Helix
shows "Not yet available" rather than estimating them.

**Inferred from the taxonomy.** Patient contact, laboratory intensity, research
intensity, commercial intensity, remote potential and travel are worked out from
what each career's own subject tags say it involves. They exist for all 716.

The second kind is genuinely useful for narrowing a list of careers, and it is
not survey data. Every screen that shows these values says so.

---

## 7. Role descriptions

### Why most careers show their family

Where an official job profile has been matched, Helix shows that profile's
description and attributes it. Where none has been matched it shows the **career
family's** description and says so, rather than generating role-specific prose.

There is no source that would fix this. The National Careers Service publishes
737 profiles and Helix uses every one that matches; NHS Health Careers publishes
around 630; ESCO's open API returns an exact title match for about 4 per cent of
the remainder and nonsense for the rest — it offered *speech and language
therapist* for Chemical Pathologist and *livestock advisor* for Medical Advisor.
Many Helix careers are simply finer-grained than anything a national service
writes a profile for.

### NHS Health Careers: linked, never copied

NHS England publishes role profiles that fit this catalogue better than anything
else available, and Helix links to 43 of them without reproducing a word.

That is a licensing decision. The National Careers Service is Crown copyright
under the Open Government Licence and may be republished with attribution. The
Health Careers terms are the opposite: they reserve all intellectual property
rights, state that the site is maintained for personal use and viewing, and
prohibit using the accompanying text for any other purpose. The same terms
explicitly permit linking, so that is what Helix does.

The pipeline enforces this structurally rather than by good intentions. It reads
only `sitemap.xml`, never requests a role page, and has no parser for one. What
it stores is a URL — not a summary, not even the page's title, so the words beside
every link come from Helix's own taxonomy. The links open in a full window
because their terms forbid framing, and the test suite fails if a link record
ever grows a field that could hold borrowed prose.



Every career has a description of its own. There are exactly two kinds, and the
data model keeps them apart so that neither can be mistaken for the other.

**`authoritative` (143 careers).** An official job profile was matched, so Helix
shows that publisher's own wording and attributes it. This text lives in the
record's `summary` field, which is the only field the attribution line and the
sources panel read.

**`taxonomy_composed` (573 careers).** No publisher has written a profile for
this exact job title, so Helix composes one from what it already records about
*that* career: its family and seniority class, its subject areas, its working
conditions (laboratory intensity, patient contact, research and commercial
work), whether statutory registration applies, and the typical entry background.
The record carries a `summary_note` saying it was composed, and it carries **no**
`source_records` — because there is no source, and attributing it to one would be
the exact dishonesty the composition is designed to avoid.

The composition is a deterministic function of those recorded fields
(`tools/market_data/describe.py`): the same career always yields the same
sentences, and the test suite re-composes published records to prove the text on
disk matches its inputs. It is closer to reading a structured record aloud than
to writing prose.

What it deliberately never says: duties, employers, day-to-day activity, or
career prospects. Helix has no source for any of that, and the tests assert those
phrasings are absent. Generating confident prose about 573 jobs from nothing
would be the fastest way to make everything else on the page untrustworthy —
composing a description strictly from recorded attributes is not that, and the
interface says which kind you are reading in both cases.

This replaced an earlier approach that showed the **career family's** description
in place of a missing role description. That was honest but not useful: fifty
careers in a family shared one paragraph, so the description could not help
anyone tell them apart. The family paragraph is still available on the career
page, folded under *About this career family*, where it reads as context rather
than as an answer.

---

## 8. How often it is checked, and how you can tell

Every record carries the date it was last checked and a date it is next due for
review.

"Last checked" means **the date the evidence was obtained**, not the date the
pipeline last ran. A record whose salary came from a National Careers Service
profile is dated the day that page was actually fetched, and re-running the
pipeline against its local cache does not move that date forward — otherwise
every record would look permanently fresh and nothing would ever be flagged for
review. Only a derived estimate carries the date of the run, because the
derivation genuinely is the thing that happened that day.

| Source | Review interval |
|---|---|
| National Careers Service | Every 90 to 180 days |
| NHS and public-sector pay | After each annual pay announcement, and at least yearly |
| ONS earnings | After each annual release |
| Derived estimates | Recomputed whenever the sources they derive from change |

A record past its review date is shown with a "due review" note. Stale data is
flagged, not hidden.

The refresh runs automatically on the third of each month and can also be
triggered by hand. It never writes to the live site directly: it opens a pull
request for a person to review, because a salary figure moving is a content
decision. Before a pull request can be opened at all, the run must show that all
every career still has a publishable salary, that the data still matches its
schema, that the careers file is untouched, and that no API key appears anywhere
in the generated output.

Any figure that has moved by more than 30 per cent is flagged for a human to look
at rather than being published silently.

---

## 9. Privacy

The end-user browser never contacts the National Careers Service, the ONS, NHS
Employers, Skills England or any salary website. Opening a career page triggers
no outside request of any kind.

All market data is read from one static file served by Helix itself. This is why
the figures never change between page views, why the application works offline
once loaded, and why using Helix reveals nothing to anybody about which careers
you looked at.

Data collection happens only in the enrichment pipeline, which runs on a build
machine, never in a browser, and never sees a CV. API keys exist only as build
secrets; the refresh fails if one ever appears in a published file.

---

## 10. Attribution

Contains public sector information licensed under the
[Open Government Licence v3.0](https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/).

Career salary and working-hours guidance: National Careers Service, Crown
copyright. Earnings statistics: Office for National Statistics, Crown copyright.

Helix does not claim ownership of Crown copyright source data.

---

## 11. Known limitations

- **NHS pay-band context is not yet populated.** Helix supports showing an
  Agenda for Change band beside the market estimate, and refuses to infer one
  from a job title. That requires the official pay scale transcribed by a person
  with its source and effective date, into
  `data/reference/nhs_pay_framework_map.json`. Until that exists, no bands are
  shown — an invented band would look official and be wrong.
- **Regional variation is not modelled.** All figures are UK-wide. The schema can
  hold regional ranges, but applying a blanket London uplift to every career
  would invent a pattern that does not exist evenly across sectors.
- **No progression forecasting.** Helix will not tell you what you might earn in
  five years. Where it shows a progression route, each step links to that
  career's own published range with its own evidence label.
- **No live vacancy data.** Helix does not know how many roles are currently
  advertised.
- **The salary and requirements evidence are independent.** A career can have a
  well-sourced salary and entirely unverified entry requirements, or the reverse.
  They are shown separately, with separate dates, so neither is read as
  vouching for the other.
- **Coverage depends on title matching.** A career with an unusual title gets a
  weaker estimate than an identical job with a common one. This is a limitation
  of matching by name, and it is why the evidence label is shown every time.

---

## 12. Salary by region

Helix has no verified career-to-SOC mapping, so it cannot publish an ONS
occupation earnings figure as a career's headline range. What it *can* say
honestly is how pay for that kind of work varies across the UK, which is a
different and much better supported claim.

**The method.** ONS ASHE Table 3 publishes median gross annual pay for full-time
employees by region for each two-digit SOC group. A region's median divided by
the UK median for the same group gives a **regional index** — 1.133 for health
professionals in London, meaning they are paid about 13% more than health
professionals nationally. Helix applies that index to the career's own UK range.
The level comes from the career's own evidence; only the regional shape is ONS's.

**Why a coarse occupation mapping is acceptable here and nowhere else.** A ratio
is far more forgiving of an approximate occupation match than a level is. The
whole-economy London index is 1.269 against 1.133 for health professionals — so
choosing the right *broad group* matters a great deal, and choosing the exact
unit group inside it barely moves the answer. Using the whole-economy figure for
a nurse would overstate London pay by about twelve percentage points; using a
neighbouring professional group would not. That is the opposite of the situation
for salary levels, where a wrong unit group produces a wrong number.

The family-to-group mapping is editorial, lives in
`data/reference/helix_family_soc_map.json`, and records a reason for every one of
the sixteen families. Seniority adjusts it: trainee and support grades move to
the associate professional group, manager and lead grades to corporate managers.

**What is never done.** A derived regional figure is capped at *indicative*
however good the UK figure was, because no source published a regional range for
that job. Regions ONS suppressed are absent from the interface rather than
back-filled with the UK figure. And there are no city-level figures: nothing in
the evidence distinguishes Manchester from Blackburn.

---

## 13. Salary by sector

This section is mostly an admission.

The question people most want answered — does industry pay more than the NHS for
this job? — has no honest answer from published data. ASHE splits earnings by
sector **or** by occupation, never both at once. There is no table of "biomedical
scientists in pharmaceutical manufacturing".

So Helix publishes exactly two things, both labelled:

1. The **public-sector pay framework** for a career, where a person has curated
   the mapping. Never inferred from a job title.
2. The **whole-economy public/private difference** from ASHE Table 25, stated as
   an all-occupations figure and explicitly not to be applied to the range above
   it.

The tempting alternative — multiplying a career's range by the average pay of
everyone employed in an industry, cleaners and directors included — would produce
an official-looking number that describes nobody.

---

## 14. Labour market signals

**Source.** ONS Faster Indicators, online job advert estimates: a weekly index of
online job adverts by advertising category, published under the Open Government
Licence and needing no credential.

**What it can say.** Whether advert volume in a category is rising, flat or
falling, and where it sits against a February 2020 baseline of 100.

**What it cannot, and therefore what Helix does not show.**

| Not shown | Why |
|---|---|
| Vacancy counts | The source is an index, not a count. Converting one to the other means inventing a total nobody published |
| Regional demand | The dataset publishes a single geography, the United Kingdom |
| Skills in adverts | Not measured by this source |
| Named employers | Not measured by this source |

**Categories, not jobs.** Signals are resolved through a career's *family*, so a
biomedical scientist's figure describes hiring across Healthcare and Social care
— a category that also covers care work. Careers in the same category read the
same number, and both the comparison table and the standout summary say so rather
than presenting one measurement twice as though it distinguished two careers.

**Age lowers the signal.** ONS last released this series in October 2024. A long,
dense weekly series is a strong measurement *of the period it covers*; if that
period ended eighteen months ago it is weak evidence about hiring today, so
strength is capped at "limited" regardless of how much data there is.

**Providers.** `tools/market_data/providers/labour_market.py` defines the
interface. Adzuna and DWP Find a Job are implemented as credentialed providers
that return nothing until a key is present in the environment, and say which key
and where to register when asked why. A stub returning plausible numbers would be
indistinguishable from a working integration in the published file. The *My data*
screen publishes the whole provider table.

**Never "no jobs".** A missing file, a failed refresh or an unmapped family all
produce "Helix has no current signal", which is a statement about Helix's
evidence and not about the job market.

## 15. Where to look next

- `docs/MARKET-DATA-AUDIT.md` — the current counts: coverage, evidence classes,
  methods, records needing review, stale records
- `docs/DATASET-AUDIT.md` — questions raised about the supplied career taxonomy
- `data/reference/helix_salary_source_registry_uk_v1.json` — the approved source
  list and the priority order
- The **My data** screen inside Helix — the same figures, in the application

---
