Take a RAMS for erecting a twelve-storey steel frame. It has a scope, a crew of six, a CPCS crane operator, four hazards with a control each, a paragraph of method, and an emergency section that says the first aider will be told and the emergency services called as required. It has no programme. It has no rescue plan for someone left hanging in a harness. It does not mention the temporary bracing design, or who checks it. It does not mention the occupied office building next door, the noise restriction, the lift plan, LOLER, or the lifting permit the site requires.
It is not a bad RAMS because it is short. It is a bad RAMS because the things missing from it are the things that kill steel erectors. And I would bet most people reading this have seen one like it get a signature, because the reviewer was checking that the document existed, was for the right trade, and had a risk assessment attached. All three were true.
That is the gap this article is about. A RAMS review is supposed to be a test. On most sites it is a receipt. I want to set out, in full, what we think a competent review tests, how much each test is worth, and how the result is scored, and I want to publish it in a form you can use on paper on Monday without ever speaking to us. The standard itself is on its own page. This is the argument for it.
What a review is legally for
RAMS is not a legal term. You will not find it in any regulation. It is how the industry chose to evidence two sets of duties that are very much legal.
The first is the Management of Health and Safety at Work Regulations 1999, regulation 3: a suitable and sufficient assessment of the risks, with the general principles of prevention in Schedule 1 telling you what "suitable" looks like: avoid the risk, then combat it at source, then collective measures before individual ones. The second is CDM 2015. Regulation 15 requires every contractor to plan, manage and monitor their own work, to apply the relevant parts of the construction phase plan, to employ nobody without the necessary skills, knowledge, training and experience, and to provide supervision, instructions and information. Regulation 13 puts the same plan, manage and monitor duty on the principal contractor for the whole phase, and adds coordination: the PC has to ensure that every contractor applies the principles of prevention and follows the plan.
A RAMS is the subcontractor's evidence under regulation 15 and the principal contractor's evidence under regulation 13. The review is the moment the PC's duty actually happens. If the review cannot say what it tested, it is thin evidence for either.
Since October 2023 there is a third layer, in England at least. The Building Regulations now carry an explicit competence regime in Part 2A. Regulation 11F says anyone carrying out building work must have the skills, knowledge, experience and behaviours to do it in a way that meets the regulations, and an organisation must have the organisational capability. Regulation 11E says the person appointing them must take all reasonable steps to satisfy themselves of that before the appointment. The guidance frameworks behind those words (BSI Flex 8670 and the PAS 8671 to 8673 series for the principal roles) are still settling, and Scotland and Wales run their own regimes, but the direction is clear. Competence is now something you have to be able to show you checked, not something you assumed because a CSCS card was photographed.
A RAMS review that produces a signature and nothing else shows none of that. A review that produces a named list of what was tested, what passed and what failed shows all of it.
Four parts, twenty tests
So here is what we test. The full standard, with weights and wording, is on the standard page. This is the walk-through.
The review splits into four parts, and the split matters because each part asks a different kind of question.
Part 1 asks whether the risk assessment is real. Two tests, and they are the heaviest in the standard at nine points each. Does every safety-significant step of the method rest on an assessed risk, with the two documents cross-referencing each other rather than stapled side by side? And does the assessment cover the hazards this activity actually presents, on this site, with controls that follow the hierarchy? That second test is where generic RAMS die. A risk assessment that lists "working at height, control: harness" for a twelve-storey frame has identified the category and skipped the work. These two carry the weight they do because they are regulation 3 of the Management Regulations in person. Everything else in the document is built on them.
Part 2 asks whether the method statement says enough. Eleven tests, fifty-six points between them, each asking the same shape of question: is this section present, specific and adequate? Scope stated with locations, boundaries, exclusions and interfaces. Programme with duration, sequence and hours. Resources, competence evidence and equipment named. Task-specific hazards, approved by someone competent. Controls that are workable: supervision, permits, inspection regimes, briefings, PPE. Emergency arrangements specific to this task and this place. Temporary works coordination, design and checks where the activity involves them. The interface with other trades and the public. Training, induction and cards. Environmental controls. Monitoring and change control.
The weights inside Part 2 are an argument, not an accident. Controls are worth eight, hazards and emergency arrangements seven each, temporary works six. Programme, environmental and monitoring are three each. That is not a statement that programme does not matter. It is a statement about what kills people when it is missing. A RAMS with no rescue plan is a RAMS under which someone can die in a harness while the site works out what to do. A RAMS with a thin programme is a RAMS that needs a conversation. The standard is weighted to say so, and every criterion maps to the duty it evidences: emergency arrangements to CDM regulations 30 and 31, temporary works to regulation 19 and BS 5975, training and induction to regulation 15(7) to (9), the public interface to regulation 13 coordination and section 3 of the 1974 Act.
Part 3 asks whether this RAMS answers the question you actually asked. Six tests, twenty-two points, and a different kind of test from everything before it. Parts 1 and 2 can be judged from the document alone; every criterion from here on is a comparison, and it needs the reviewer to hold the requirement in one hand and the submission in the other. The standard makes that split explicit: a review run without the requirement can only run the thirteen document-only criteria, and has to say so. This is also the part most paper reviews skip entirely. When the principal contractor issued the requirement for this activity, it named a risk category, cited the legislation that applies, said whether a permit is needed, referenced any temporary works, described the activity, and set a risk level. Part 3 checks each of those against the submission. Does it engage the risk category? Does it reflect each piece of legislation cited, not in general but specifically? Does it reference raising and closing the permit? Does it address the temporary works you referenced, not the ones it chose to mention? Does its scope match the activity described, neither falling short nor overrunning? Are the controls proportionate to the risk level you set, rather than to a generic baseline?
Part 3 is where the steel-frame RAMS fails hardest, because the requirement said tower crane operations, lifting, LOLER, BS 7121, a lifting permit and a crane foundation, and the submission answered none of it. It is also the part that makes the review site-specific in the legal sense. Regulation 13 coordination is exactly this: the PC defined what this activity needs and then checked that what came back met it.
Part 4 asks whether it meets your own rules. One test, four points, for the organisation's own requirements: the client's standards, the framework conditions, the PC's site rules that go beyond the regulations. It only applies if you have issued such rules. If you have not, the criterion is marked as not applicable and comes out of the score entirely, so nobody is marked down for a rule that was never set. And the standard is explicit that the reviewer must not invent organisation rules to fill the gap.
Why it is scored, and how
The obvious objection to scoring a RAMS is that safety is not a percentage. I agree, and that is why the score is designed the way it is.
Every criterion gets a verdict (met, partially met, not met, or does not apply) and a severity (critical, major, minor, or none). The verdict and severity are judgements, made by a reviewer or, in our case, by an AI reviewer working from the document. The score is never a judgement. It is arithmetic over the verdicts. A criterion's deduction is its weight multiplied by how far short it fell (not met counts in full, partially met counts half) multiplied by how much it matters (critical in full, major at six tenths, minor at a quarter). Criteria that do not apply come out of the denominator. Because no deduction can exceed the criterion's own weight, the result sits inside 0 to 100 by construction.
That gives you a number where every point lost is traceable to a named criterion. If a subcontractor challenges a score of 61, the answer is not "the reviewer felt it was about a 61". The answer is a table: criterion 8, emergency arrangements, not met, critical, seven points; criterion 16, permit process, not met, major, 1.8 points; and so on to the total.
Then there are three caps, and the caps are the most important design decision in the standard.
Weighted arithmetic on its own has a flaw. A RAMS that is excellent everywhere and has no rescue plan scores in the high eighties or nineties, because one criterion out of twenty can only take so many points. That is the wrong answer. No rescue plan on work at height is not a ninety. So the standard says: if any criterion is not met at critical severity, the score is capped at 49, the bottom band, whatever else the document does well. If any criterion is partially met at critical severity, the cap is 69. And if any criterion is not met or partially met at major severity, the cap is 89.
One thing to hold onto when reading the bands: a band describes the document as a whole, never an individual finding. One major finding caps the document at 89, into the band called gaps to close, precisely because the top band is defined by the absence of material gaps.
The last cap exists because of a specific failure we found when testing the arithmetic. A submission with two major findings on low-weight criteria scored 96, which put it in the top band next to a written review that listed material gaps. The top band is defined as "no material gaps; any findings are minor". A major finding contradicts that definition, so a major finding has to put the score below 90, however small the criterion's weight. The score and the verdict now cannot disagree.
Two more rules, both about refusing to produce a number. A score built on a partial assessment is not a score: if any criterion has no verdict, the review reports the score as unavailable and names the criteria that were missed. And a score of zero is never used to mean "we could not read the output". Zero is a real score, reserved for a submission that fails every criterion critically. When the review could not be read, the score is null with a reason. The reason that matters is that a silent zero and a real zero look identical on a dashboard, and only one of them should have someone reaching for the phone. This is the same principle as the content gate in the hallucination article: an assessment of nothing must be impossible.
The steel frame, scored
Run the RAMS from the opening through the standard. Scope stated: met. Risk assessment generic: partially met, major. No programme: not met, major. No competence evidence: partially met, major. Hazard register generic, structural instability during erection unaddressed: not met, critical. No briefings: partially met, major. No rescue plan: not met, critical. No temporary works design, no TWC: not met, critical. Adjacent occupied building unaddressed: not met, major. No environmental controls: not met, major. Monitoring thin: partially met, minor. Lifting risk category unaddressed: not met, critical. LOLER and BS 7121 unaddressed: not met, critical. Permit not referenced: not met, major. Crane foundation temporary works partly addressed: partially met, major. Scope broadly matches: partially met, minor. Controls not proportionate to a high risk level: not met, major. No organisation rules supplied: does not apply.
The deductions come to 47.2 points against 96 applicable, which is 50.9 out of 100 before caps. Then five criteria failed at critical severity, so the cap bites: 49. Bottom band: serious gaps. Back to the subcontractor.
Look at what the arithmetic did and did not do there. It did not decide the RAMS had to go back; the five critical verdicts did that, and they are each defensible on their own. What the arithmetic did was make the result reproducible. Two reviewers with the same verdicts get the same number. And when the subcontractor's revision comes back with a TWC appointed, a rescue plan, a lift plan and the adjacent building addressed, the same twenty tests show exactly which points came back.
The two scope tests
Someone will notice that scope appears twice, as criterion 3 in Part 2 and criterion 18 in Part 3, and will reasonably ask whether that is a double count. It was, nearly, in the first version. Up to v1.0.0 the two were worded so similarly that a scope mismatch could be penalised twice. In v1.1.0 we reworded them rather than merging them, because they test different things and the remedies are different.
Criterion 3 asks the Part 2 question: does the method statement state a scope at all, with locations, boundaries, exclusions and interfaces, specifically enough that the rest of the document can be audited against it? A vague scope means rewrite the method statement. Criterion 18 asks the Part 3 question: does that scope conform to the activity the requirement describes? A mismatched scope means you have been sent the wrong RAMS for this work. Most audit checklists separate contents from conformity for the same reason, and the published wording now says explicitly which test each one is and which it is not.
I mention it because a standard that has been argued over is worth more than one that has not, and because if you adopt this on paper, that is the line your own reviewers will ask about first.
What this standard does not do
It does not test whether the controls will work on the day. A RAMS can pass all twenty tests and be ignored by the gang at seven on Monday morning. That is supervision and monitoring, criterion 13 only tells you whether the document describes it.
It does not replace judgement on severity. Whether a missing briefing is major or minor on this job is a call, and the standard gives the reviewer the arithmetic to make that call count, not a rule that makes it for them. An experienced H&S manager will set severities differently from a junior, and should.
It is not a legal opinion. The duty mapping on the standard page says which regulation each criterion evidences. It does not say that passing the test discharges the duty. That is for you, your advisers and, in the worst case, a court.
And it is not a British Standard. It is a published review standard, version-controlled and open, offered for adoption and for criticism. We will change it when someone shows us it is wrong, and the changelog will say who did.
Adopt it, then, if you want, automate it
The standard page carries the twenty criteria with their weights, the duty each one evidences, the verdict and severity factors, the caps, the bands and the refusal rules, plus the worked example above. There is a one-page checklist you can print. It is licensed for free use with attribution. If your reviewers use it on paper and never buy anything from us, it has done its job: there will be one more site where the RAMS review is a test and not a receipt.
There is also a free version of the review you can run right now at planops.ai/rams-review: upload a RAMS, no account, and it is read against the thirteen document-only criteria, Parts 1 and 2 of this standard. It cannot run Parts 3 and 4, because those compare the submission with the requirement you issued and your own rules, and a public page has neither. The result says so rather than pretending otherwise: the standard treats a criterion with no comparison material as not applicable, never as a failure and never as a guess.
Inside PlanOps, where the requirement and your organisation rules exist on the project, all twenty criteria run on each RAMS that arrives, with every verdict written into the report, every point traceable to its criterion, and the AI never asked for a number. The plain-English tour of how that fits into the rest of the job is at planops.ai/workflows.
Either way, the question to ask of any RAMS review, whoever does it, is the same. What did it test? If the answer is a list, you have a review. If the answer is a signature, you have a receipt.
We write about what we are learning building AI for UK construction, the failures included, in the Construction AI Brief. If this was useful, you can get it monthly.
Ian Yeo is the founder of PlanOps, an AI-native planning operations platform for UK construction.
Read next: How we stop AI making things up and Why can't we just use a chatbot?
