Ronan Blake
Free · No signup

275 questions UK interviewers actually ask.

Seven domains, three levels, with a full answer and — more useful — a note on what the interviewer is really assessing. Search it, filter it, or work through a domain before an interview.

Security Operations / SOC

52
What does a SOC actually do, day to day?entrytechnicalCommonly asked

It monitors an organisation’s environment for signs of attack, investigates what the tooling surfaces, and responds to or escalates what turns out to be real. In practice most of the day is triage: working a queue of alerts, most of which are benign, and deciding quickly which are not — a phishing report here, an EDR alert for a suspicious script there, a login from an unexpected country. Alongside that queue sits the proactive work that stops the queue getting worse: tuning noisy detections, hunting for activity that never alerted, onboarding new log sources, and writing the rules that generate tomorrow’s alerts. A SOC also has a rhythm the job adverts rarely mention — shift handovers, on-call rotations, and the discipline of documenting what you did so the next analyst can pick it up. The romantic version is constant incident response; the honest version is disciplined, repetitive investigation punctuated by occasional genuine incidents, and the analysts who thrive are the ones who stay sharp through the repetition rather than the ones waiting for the dramatic day.

What they are assessing

Whether you have a realistic picture of the work rather than one built from television.

What is the difference between Tier 1, Tier 2 and Tier 3 in a SOC?entrytechnicalCommonly asked

Tier 1 handles initial triage: working the alert queue, applying documented playbooks, closing the benign and escalating anything that is not. Tier 2 takes those escalations and investigates properly — deeper analysis, correlating across log sources, driving containment on real incidents. Tier 3 covers the specialist work: threat hunting, detection engineering, forensics, and the hardest incidents that Tier 2 cannot resolve alone. The progression describes how responsibility changes in kind, not just difficulty: Tier 1 applies process, Tier 2 exercises judgement within it, and Tier 3 defines what the process should be. Two honest caveats worth raising, because they show you understand the real landscape. First, the model is not universal — flatter, blended teams are increasingly common, and many modern SOCs deliberately push investigation skills down to Tier 1 rather than gate-keeping them. Second, Tier 1 is a starting point, not a ceiling or a dead end: it is where you build the pattern-recognition that everything above depends on, and the analysts who move up fastest are the ones who treat triage as training rather than as a queue to survive.

What they are assessing

Whether you understand the career progression you are entering, and that Tier 1 is a starting point rather than a ceiling.

What is a playbook, and what are its limits?entrytechnical

A documented procedure for handling a specific alert or incident type — the steps to take, what to check, the decision points, and when to escalate. Their real value is consistency and speed: a playbook lets a new analyst handle a phishing report or a malware alert competently on day two rather than month six, and it means the same alert gets handled the same way regardless of who catches it or how tired they are at 3am. The limit is fundamental: a playbook encodes what was anticipated when it was written. An attacker doing something genuinely novel produces an alert whose playbook does not quite fit, and an analyst following the steps mechanically will reach the “close” action and move on — the playbook actively steers them past the thing that mattered. This is why good playbooks are written as a floor for judgement rather than a substitute for it: they handle the routine so attention is freed for the anomalous. The analysts worth promoting are precisely the ones who notice when the playbook and the evidence disagree, and stop rather than tick the last box. A strong answer names the maintenance problem too — a playbook nobody has reviewed since the tooling changed is quietly wrong, and stale playbooks are worse than none because they carry false authority.

What they are assessing

Whether you see process as support for judgement rather than a replacement for it.

What is the difference between a SIEM and a SOAR platform?midtechnical

A SIEM aggregates and analyses log data to detect and surface suspicious activity — it is the detection and investigation layer. SOAR sits alongside it and automates the response: orchestrating actions across tools, running automated enrichment and containment, and managing case workflow. In simple terms, the SIEM tells you something happened and SOAR does something about it without a human doing every step by hand. Most mature SOCs run both, with SOAR handling the repetitive enrichment that used to consume analyst time — pulling reputation data, checking asset ownership, isolating a host on confirmation.

What they are assessing

Whether you understand the modern SOC tool stack rather than just the SIEM.

What is EDR, and what does it give you that logs alone do not?entrytechnicalCommonly asked

Endpoint Detection and Response is an agent on the host that records detailed endpoint activity — process execution, file changes, registry modifications, network connections — and can act on it: isolating the machine, killing processes, collecting artefacts. What it adds over log collection is depth and causality: instead of knowing a connection happened, you can see which process made it, what launched that process, and what it touched afterwards. That process lineage is what turns an alert into an understood incident, and it is why the endpoint has become the primary source of truth as network traffic has encrypted.

What they are assessing

Whether you understand why endpoint visibility became central rather than just naming the product category.

What is XDR, and is it meaningfully different from EDR?midtechnical

Extended Detection and Response broadens EDR beyond the endpoint, correlating signals from endpoint, identity, email, cloud and network in one platform so a single attack chain can be followed across all of them rather than landing as separate, disconnected alerts in separate consoles. The value is concrete: a phishing email, the endpoint that then ran a suspicious script, and the identity that subsequently logged in from a new location are one story, and XDR’s pitch is to stitch that story together automatically instead of relying on an analyst to connect three tools at 2am. The honest answer on whether it is *meaningfully* different is where you show maturity: the correlation is genuinely useful, but “XDR” is one of the most heavily marketed terms in the industry, and what one vendor calls XDR another calls a SIEM with good integrations, or simply EDR with a wider telemetry net. The distinction that matters is not the acronym but two practical questions — what does this specific product actually correlate, and does it reduce the number of consoles your analysts have to watch? If it genuinely collapses five tools into one coherent timeline, it earns the name; if it is a dashboard bolted over the same silos, it does not, whatever the datasheet says.

What they are assessing

Whether you can discuss vendor categories with appropriate scepticism.

Walk me through triaging a malware alert from EDR.entrytechnicalCommonly asked

First, what fired and on what: which detection, which host, which user, and how confident the tooling is. Then the process context — what executed, what launched it, the full lineage, and whether the path and signing look legitimate. Then whether it actually ran or was blocked, because that changes urgency completely. Then blast radius: did it make network connections, touch files, create persistence, or spread. In parallel I would check whether the same detection has fired elsewhere, because one host is an incident and twenty is a campaign. Contain if confirmed, document throughout, escalate on anything beyond my authority.

What they are assessing

Structured investigation with process lineage and blast radius, not just reading the alert text.

An alert fires for PowerShell with an encoded command. How do you investigate?midtechnicalCommonly asked

Encoded PowerShell is common in attacks and also in legitimate administration, so the first job is decoding rather than concluding. Decode the base64 and read what it actually does. Then context: which user and host, what launched PowerShell — a parent process of Office or a browser is far more suspicious than a management tool — and whether it runs on a schedule. Check whether the decoded content reaches out to a network location, downloads anything, or establishes persistence. Legitimate encoded commands tend to be repeatable, documented and launched by known tooling; malicious ones usually appear once, from an unexpected parent.

What they are assessing

Whether you decode and reason rather than treating encoding itself as proof of malice.

What parent-child process relationships would immediately concern you?midtechnical

The ones that should not happen in normal operation. Office applications spawning command interpreters — Word launching PowerShell or cmd — is a classic macro-attack signature. Browsers spawning shells. Web server processes launching interpreters, which suggests a web shell. Anything spawning from a temporary directory or a user’s download folder. Unusual chains involving system utilities called by user applications. None of these is proof on its own, because legitimate software does strange things, but each is worth a look, and detections built on abnormal lineage tend to be more durable than ones built on file signatures.

What they are assessing

Whether you can name specific suspicious patterns rather than talking generally about anomalies.

What is "living off the land", and why does it complicate detection?midtechnical

Attackers using legitimate tools already present on the system — PowerShell, WMI, certutil, scheduled tasks, remote management utilities — rather than bringing their own malware. It complicates detection because there is no malicious file to find and no signature to match: the binaries are signed, expected, and used constantly by administrators. Detection therefore shifts from what ran to how and by whom — unusual parent processes, unexpected command-line arguments, an account that has never used a tool suddenly using it, activity at odd hours. It is the main reason behavioural detection has displaced signature matching.

What they are assessing

Whether you understand why modern detection is behavioural rather than signature-based.

How would you detect credential dumping?midtechnical

The classic signal is unexpected access to the memory of the process holding credentials — on Windows, suspicious handle requests to LSASS, which almost nothing legitimate does. Beyond that: known tooling signatures, though attackers rename and modify them; unusual access to registry hives holding credential material; volume shadow copy creation by an unexpected process; and, on the outcome side, an account suddenly authenticating from somewhere new shortly after suspicious activity on a host. The strongest detections combine the memory-access signal with process context, and it is worth knowing that legitimate security tooling also touches LSASS, so tuning matters.

What they are assessing

Specific technical knowledge of a common attack step, plus awareness of false-positive sources.

What is Pass-the-Hash, and how would you spot it?midtechnical

An attack where the attacker authenticates using a stolen password hash without ever knowing the plaintext password, exploiting the fact that some authentication protocols accept the hash as proof. Detection is behavioural rather than signature-based: look for authentication patterns that do not fit the account — logons to systems the user never touches, NTLM authentication where Kerberos would be expected, the same account authenticating from multiple hosts in quick succession, and administrative logons from workstations rather than management systems. The underlying defence is limiting where privileged credentials are exposed in the first place.

What they are assessing

Understanding of a foundational lateral-movement technique and its behavioural signals.

What is Kerberoasting, and what does it look like in logs?midtechnical

An attack against Active Directory where an authenticated user requests service tickets for accounts running services, then cracks those tickets offline to recover the service account password — no elevated access needed to make the request, which is what makes it attractive. In logs it appears as service ticket requests, particularly a single account requesting many tickets in a short period, and requests using weaker encryption which is easier to crack. Defences are strong passwords on service accounts, managed service accounts where possible, and alerting on unusual volumes of ticket requests from one principal.

What they are assessing

Depth in Active Directory attacks, which dominate real UK enterprise incidents.

How would you investigate a suspected web shell?midtechnical

Start from the web server process: a web shell shows up as the web server spawning a command interpreter or other unexpected child process, which almost never happens legitimately. Then the file system — recently created or modified files in web-accessible directories, especially with unusual extensions or timestamps inconsistent with a deployment. Then the web logs, looking for requests to unfamiliar paths, often with parameters carrying commands, and frequently from a single source. Then outbound connections from the server. Containment is usually isolating the host and preserving the file before removal, because the shell is evidence of how they got in.

What they are assessing

Whether you can investigate a server-side compromise rather than only endpoint alerts.

A user reports their machine is slow. When does that become a security investigation?entryscenario

Usually it does not — most slow machines are slow for ordinary reasons, and treating every performance complaint as an incident wastes the goodwill you need when it matters. It becomes security-relevant when the performance issue has a suspicious shape: sustained high CPU from an unfamiliar process, which can indicate cryptomining; unusual network volume; the slowness starting immediately after the user opened something; or the machine belonging to someone likely to be targeted. The practical approach is a quick look at running processes and network connections, which costs minutes and either closes it or escalates it.

What they are assessing

Proportionate judgement, and awareness that user reports are a real detection source.

How would you detect cryptomining on corporate systems?entrytechnical

The behavioural signals are distinctive: sustained high CPU or GPU usage that does not match the user’s work, often continuing outside working hours; connections to known mining pool infrastructure or unusual persistent outbound connections; and process names that are unfamiliar or masquerading as system binaries. On servers it often shows as a performance complaint before a security alert. Detection can lean on threat intelligence for pool domains, but the durable signal is the resource pattern, because pool infrastructure changes. It is worth knowing that mining is frequently a symptom rather than the whole problem — the interesting question is how it arrived.

What they are assessing

Whether you look past the obvious symptom to how the system was compromised.

What makes a good detection rule?midtechnicalCommonly asked

It catches the behaviour reliably, produces few false positives at your scale, and gives the analyst enough context to act. Beyond that, the best rules target behaviour rather than artefacts — detecting the technique rather than a specific file hash or IP, because those change and the technique persists. A good rule is also documented: what it detects, why it matters, what the analyst should do, and what known-benign sources exist. And it is testable — you should be able to trigger it deliberately and confirm it fires, because a rule nobody has ever seen work is a rule you are trusting on faith.

What they are assessing

Whether you think about detections as engineered products with maintenance costs.

What is the Pyramid of Pain, and why does it matter for detection?midtechnical

A model ranking indicator types by how much difficulty detecting them causes an attacker. At the bottom, hashes and IP addresses — trivial for an attacker to change, so detections built on them are brittle. Higher up, domains and network artefacts. At the top, tools and tactics, techniques and procedures — changing those means changing how the attacker works, which is expensive. The practical implication is where to invest: hash-based blocking is cheap and worth having, but the detections that actually cost an adversary something are behavioural. It also explains why threat-intel feeds of IPs deliver less value than expected.

What they are assessing

Whether you understand detection strategy rather than only detection mechanics.

How would you tune a detection rule producing too many false positives?midtechnicalCommonly asked

First understand why it is firing: pull a sample of recent hits and find what the benign ones have in common — a particular process, service account, subnet, or time window. Then decide whether to narrow the rule, add exclusions, or change the logic to require additional conditions. Narrowing is better than excluding where possible, because exclusions accumulate into holes. Document every change with the reason, because an undocumented exclusion is indistinguishable from a gap two years later. Then measure: is the volume manageable now, and have you tested that the rule still fires on the real behaviour?

What they are assessing

Methodical tuning with awareness that exclusions create long-term risk.

How would you use MITRE ATT&CK to find gaps in your detection coverage?seniortechnical

Map the detections you actually have to the techniques they cover, which usually reveals that coverage is concentrated in a few tactics — typically execution and command-and-control — and thin elsewhere, often in persistence and defence evasion. Then prioritise the gaps by relevance rather than trying to cover the matrix: which techniques do the threat actors targeting your sector actually use, and which map to your crown jewels. The honest caveat is that mapping tends to overstate coverage, because a detection that catches one narrow implementation of a technique gets marked as covering the whole thing.

What they are assessing

Practical use of the framework, including scepticism about coverage claims.

What is detection as code, and why does it matter?seniortechnical

Treating detection rules as software: stored in version control, peer-reviewed before deployment, tested automatically, and released through a pipeline rather than edited directly in the console. It matters because detections are production logic with real consequences, and the traditional approach — an analyst editing a rule live, with no history and no review — produces exactly the problems you would expect: nobody knows who changed what or why, changes break things silently, and knowledge lives in individuals. Version control also makes it possible to answer the question that matters after an incident: what were we detecting on that date?

What they are assessing

Whether you bring engineering discipline to detection work.

How would you test that a detection actually works?seniortechnical

Trigger the behaviour deliberately in a controlled way and confirm the alert fires with the context the analyst needs — not just that *something* fired, but that it fired with enough detail to act on. Atomic testing frameworks exist for exactly this: Atomic Red Team, for instance, executes individual MITRE ATT&CK techniques safely so you can validate whether your coverage actually catches them, and purple-team exercises do the same thing collaboratively at greater depth. But the single test is the easy part. The real discipline is repetition, because detections decay silently: a log source changes format, an agent stops reporting, someone adds a well-meaning exclusion to cut noise, a field gets renamed upstream — and the rule that worked at build time now matches nothing, with no error and no alert to tell you. This is the trap. A detection that fired correctly the day it was written and has never been re-tested since is a control that *exists* rather than one that *works*, and the gap between those two is where breaches live. A strong answer therefore frames detection testing as continuous validation — ideally automated and scheduled — rather than a one-off sign-off, and treats “when did we last prove this fires?” as a question every important detection should be able to answer.

What they are assessing

Verification instinct, and awareness that detections decay.

How would you know if a critical log source stopped reporting?midtechnical

You would need a detection for the *absence* of data — a “dead man’s switch” — which is the control most SOCs forget to build precisely because nothing about a missing log source announces itself. The mechanism is to monitor ingestion volume per source and alert when a source drops below its normal baseline or stops entirely. Doing that well means baselining each source’s expected pattern, including its rhythm: a domain controller logs steadily around the clock, whereas a payroll system might legitimately go quiet at weekends, so a naive “no events in an hour” rule would either miss real outages or drown you in false positives. It matters because a silent log source is worse than a noisy one. A noisy source generates alerts you can triage; a silent one generates nothing, the dashboard stays green, and the team assumes a coverage they no longer have — a gap that can persist for weeks. Worse, log outages are not always accidental: disabling or starving logging is a recognised attacker technique (it appears in MITRE ATT&CK as Impair Defenses), so an outage frequently coincides with exactly the activity you most wanted to see. A strong answer names the second-order version too: it is not enough to alert on a source going dark — someone has to *own* the response, because an unactioned “source X stopped reporting” ticket sitting in a queue for a fortnight is functionally the same as having no detection at all. Silence is the most dangerous state in a SOC.

What they are assessing

Whether you think about detection failure modes rather than only detections.

What log sources would you prioritise if you could only have five?midtechnicalCommonly asked

For most organisations: identity provider authentication logs, because identity is where modern attacks concentrate; endpoint telemetry from EDR, because that is where execution is visible; DNS, because almost every attack chain touches name resolution and it is remarkably cheap to collect; firewall or proxy logs for egress visibility; and domain controller logs in a Windows estate, because Active Directory is the crown jewel. The reasoning matters more than the list — you are choosing sources that cover the stages an attacker must pass through, rather than collecting whatever is easiest to enable. NCSC’s logging and protective monitoring guidance is a useful sanity check on the list, and is what a UK public-sector panel will expect you to have read.

What they are assessing

Prioritisation reasoning against attacker behaviour rather than a memorised list.

Why is DNS logging disproportionately valuable?midtechnical

Because almost everything touches DNS before it does anything else, and it stays legible when other traffic does not. A near-complete list of things that show up in DNS first: malware resolving command-and-control domains, data exfiltration over DNS tunnelling, users reaching phishing infrastructure, beaconing to newly registered domains, and lateral movement resolving internal hostnames. It is also cheap — DNS logs are tiny and retainable for long periods compared with full packet capture, so you can keep months of history for the cost of keeping days of everything else, which matters enormously when you discover an intrusion that began long ago and need to trace it backwards. The deeper reason it punches above its weight is the encryption trend: as more traffic moves to TLS, the payload goes dark, but the DNS lookup that *precedes* the connection often remains visible, so DNS is one of the last places where intent is legible before the wire goes opaque. The honest senior-level caveat is that this is eroding — encrypted DNS (DoH and DoT) is closing that window, and an attacker using DoH to a reputable resolver bypasses your DNS logging entirely — so a strong answer pairs “DNS is disproportionately valuable” with “and part of the job now is making sure endpoints are forced through DNS you can actually see.”

What they are assessing

Understanding of why certain data sources punch above their weight.

How would you detect DNS tunnelling?seniortechnical

Look for the statistical fingerprint rather than any single query, because no individual lookup looks malicious. The tells: unusually long or high-entropy subdomain strings (data encoded into the query name), a very high volume of queries to a single parent domain, unusual record types used consistently (TXT and NULL records carry more data than A records), and query timing too regular to be human. The classic pattern is a single host generating thousands of unique subdomains under one parent domain in a short window — no human or normal application behaves that way. In practice you detect it by baselining per-domain query volume and subdomain entropy and alerting on outliers, rather than trying to write a signature for “tunnelling.” The complication, and the thing an interviewer wants you to raise unprompted, is false positives: several legitimate services genuinely use DNS this way — some antivirus and security products encode lookups into DNS, and certain CDNs and telemetry systems generate high-entropy subdomains by design. If you alert without baselining those out first, the rule produces noise, the analysts learn to ignore it, and it gets tuned into oblivion within a week — so the detection is only as good as the allow-listing and environmental knowledge behind it.

What they are assessing

Detection thinking based on patterns rather than indicators, plus false-positive awareness.

How would you approach log retention decisions?seniortechnical

Balance three pressures: detection needs, investigation needs, and cost. Detection mostly needs recent data — rules run on what is arriving now. Investigation needs history, and the uncomfortable fact is that intrusions are frequently discovered months after they begin, so thirty days of retention means the beginning of the incident is gone. Regulatory and contractual requirements set a floor in some sectors. The common compromise is tiered: recent data hot and searchable, older data in cheaper storage that is slower to query but still available. What matters is deciding deliberately rather than defaulting to whatever the licence includes. In UK regulated sectors, check whether specific retention floors apply before optimising — financial services and public-sector contracts frequently set them, and discovering that after the fact is expensive.

What they are assessing

Whether you connect retention to the reality of dwell time and cost.

How would you reduce SIEM ingestion costs without losing visibility?seniortechnical

Start by finding out what is actually being used: a substantial proportion of ingested data typically never appears in a detection or an investigation. Filter noise at source — verbose debug logging, high-volume low-value events — rather than paying to store and index it. Route bulk low-value data to cheaper storage that can still be searched when needed, keeping the SIEM for what detections run against. And challenge the assumption that everything must be centralised. The discipline is deciding by value rather than by volume, and reviewing periodically, because ingestion grows silently as new systems come online.

What they are assessing

Commercial awareness alongside technical judgement — increasingly asked as SIEM costs rise.

What is dwell time, and why does it matter?midtechnicalCommonly asked

The period between an attacker gaining access and being detected. It matters because almost everything an attacker achieves — reconnaissance, privilege escalation, lateral movement, staging data — happens during it, so reducing dwell time reduces impact even when you cannot prevent the initial compromise. It is also a more honest measure of a SOC than alert volume: a team closing thousands of alerts while missing an intrusion for four months is busy rather than effective. Industry dwell-time figures vary widely by source and sector, so quote the concept confidently and specific numbers carefully.

What they are assessing

Whether you measure detection capability by outcome rather than activity.

How would you measure whether a SOC is effective?seniortechnical

Not by alert volume, which measures how noisy the tooling is. The meaningful measures are outcome-shaped: time to detect and time to respond for real incidents; how incidents were discovered, because a high proportion found by external parties is a poor sign; detection coverage against the techniques relevant to your threat model; and the false-positive rate, because it determines whether analysts can sustain the work. Alongside those, one honest qualitative measure: when the team tests a detection deliberately, does it fire? Metrics that reward activity produce busy teams; metrics that reward outcomes produce effective ones. If you are interviewing in UK financial services, expect the operational-resilience framing too: can you evidence that monitoring supports the firm’s important business services staying inside their impact tolerances.

What they are assessing

Metric literacy, and resistance to vanity numbers.

What is the difference between containment, eradication and recovery?entrytechnicalCommonly asked

Containment stops the incident spreading or continuing — isolating hosts, disabling accounts, blocking connections — and is urgent. Eradication removes the attacker’s presence: the malware, the persistence mechanisms, the accounts and access they created. Recovery restores normal operation, which may mean rebuilding systems and confirming they are clean before returning them to service. The order matters and the common failure is recovering before eradicating — restoring a system that still holds the attacker’s persistence, so the incident restarts. The other common failure is destroying evidence during containment that you needed for eradication. One UK-specific point worth carrying: if personal data is involved, the ICO clock starts at awareness of a likely breach, not at the end of your investigation, so the reporting decision runs in parallel with the technical work rather than after it.

What they are assessing

Understanding of incident phases and the failure modes between them.

How would you contain a compromised host while preserving evidence?midtechnicalCommonly asked

Network isolation through the EDR agent, rather than pulling the cable or powering off. Isolation cuts the host’s ability to reach anything else — stopping lateral movement and severing the attacker’s access — while leaving the machine running and the agent connected, so you can keep investigating remotely and, crucially, so volatile evidence survives. This is the heart of the question: *how* you contain changes *what evidence you keep*. Powering off destroys everything held only in memory — running processes, live network connections, injected code that never touched disk, encryption keys — which is frequently exactly the material that reveals what the attacker was doing. Pulling the network cable stops the bleeding but also kills your remote visibility and tempts someone to walk over and start poking at the console, contaminating the timeline. So the order is: isolate at the network level via the agent; if memory capture is warranted, do it before any other action that might disturb state; then investigate. Throughout, document what you did and precisely when, because during analysis you must be able to separate the attacker’s actions from your own — an undocumented responder action can look identical to adversary activity and send an investigation down the wrong path for hours.

What they are assessing

Whether you know that containment method affects evidence, and choose accordingly.

What is order of volatility?midtechnical

The principle that evidence should be collected from most to least perishable, because some sources vanish the moment you touch the machine and others persist for months. The rough order, most volatile first: CPU registers and cache; the contents of RAM — running processes, live network connections, decrypted data, injected code that never touches disk; the routing and ARP tables; temporary filesystems; then disk; then remote logging; and finally archived or physical backups. The practical consequence is blunt: rebooting or shutting down a suspicious machine destroys the most valuable evidence first — memory-resident malware, the attacker’s live connections, encryption keys held only in RAM — none of which survive a power cycle. The instinct to “turn it off and on again,” or to reimage a box to get the user working, is exactly the instinct that ruins an investigation. The interviewer is usually not testing whether you can perform forensic acquisition yourself — most SOC analysts never will. They are testing whether, in the first ten minutes of a suspected compromise, you know what *not* to destroy: don’t reboot, don’t shut down, don’t reimage; isolate the host at the network level so it stops talking to the attacker while preserving its state, and escalate to whoever will capture memory before anyone touches the disk. Knowing the order is really about keeping your options open for the person who comes after you.

What they are assessing

Forensic awareness sufficient to avoid ruining an investigation.

What is chain of custody, and when does it matter in a SOC?midtechnical

A documented record of who handled evidence, when, and what they did with it — establishing that the evidence has not been altered. It matters whenever the incident might end up somewhere consequential: a disciplinary process, a police investigation, litigation, or a regulatory examination. The practical SOC version is disciplined documentation — hashing collected artefacts, recording timestamps and handlers, storing copies securely, working on copies rather than originals. You will not always know at the outset whether an incident will become a legal matter, which is the argument for handling evidence properly by default. In the UK, the ACPO principles for digital evidence remain the common reference point, and being able to name them signals you have thought about this beyond the SOC.

What they are assessing

Whether you understand that incidents sometimes have legal afterlives.

A lessons-learned review is scheduled after an incident. What makes it useful?midcompetency

Blamelessness, first and above everything — if people expect to be punished, the account you get will be curated, the person who clicked the link goes quiet, and the real root cause stays buried. Psychological safety is not a nicety here; it is the precondition for getting true information at all. With that in place, focus on the system rather than the individual: not *who* clicked, but *why clicking had that consequence* — why the email reached the inbox, why the attachment executed, why the endpoint let it, why nothing detected the follow-on activity. Useful reviews produce specific, owned actions with named owners and due dates — “Jaz to add a detection for this persistence technique by the 14th” — not general resolutions to “be more vigilant,” which change nothing. And they examine detection as rigorously as response: not just how well we handled it once we knew, but why we did not see it sooner and what would have caught it earlier. The real test of the review is not the meeting; it is whether anything measurably changed ninety days later. A review whose actions are all still open a quarter on was theatre, and experienced teams track lessons-learned actions to closure exactly as they would incident tickets.

What they are assessing

Whether you understand incidents as improvement opportunities and know what spoils them.

What is threat hunting, and how does it differ from monitoring?seniortechnicalCommonly asked

Monitoring is reactive: the tooling raises something and an analyst investigates it. Hunting is proactive and hypothesis-driven: an analyst forms a specific idea about how an attacker might be operating *undetected* in the estate, then goes looking in the data for evidence of it, with no alert prompting them. The distinction people miss is that hunting is not “looking around in the logs for anything odd” — that is browsing, and it does not scale or repeat. A hunt starts from a testable proposition. It exists because detections, by definition, only catch what they were built to catch, and a capable adversary studies to operate in the gaps between them; hunting is how you go looking in those gaps deliberately rather than waiting to get lucky. A crucial point that separates people who have actually hunted from people who have read about it: a hunt does not fail if it finds nothing. A well-scoped hunt that comes back clean is genuine evidence of absence for that hypothesis, which has real value — and almost every hunt, successful or not, produces a durable by-product: a new detection, a data-quality gap discovered, a blind spot mapped. That conversion of one-off hunt into permanent automated coverage is arguably hunting’s main long-term payoff, not the occasional caught intruder.

What they are assessing

Whether you understand hunting as hypothesis-driven rather than as looking around.

How would you structure a threat hunt?seniortechnical

Start with a hypothesis specific enough to test — not “look for attackers,” but “if an adversary were using scheduled tasks for persistence in our estate, what would that look like in our data, and where would it be?” The specificity is the whole game: a vague hunt produces a vague result you cannot act on or repeat. From the hypothesis, derive what evidence would exist if it were true and which data source would hold it, then confirm you actually collect that data — discovering mid-hunt that the relevant logging does not exist is itself a valuable finding. Then query, and expect the overwhelming majority of what you find to be legitimate; the real work is separating unusual-but-benign from genuinely suspicious, which is impossible without knowing your own environment, so hunting rewards familiarity more than raw tool skill. Frameworks like MITRE ATT&CK help by giving you a structured menu of adversary techniques to hunt against rather than starting from a blank page. Whatever the outcome, document the hypothesis, the data examined, and the result — a clean hunt is only evidence if someone recorded what was checked — and convert anything durable into a scheduled detection, so a hunt that proved valuable never has to be run by hand again.

What they are assessing

Structured method rather than unstructured curiosity.

How would you use threat intelligence in a SOC?midtechnical

Three ways, in ascending order of value. Tactically, as indicators fed into detection and blocking — useful but perishable, and the lowest-value use. Operationally, to understand which techniques the actors targeting your sector actually use, which then drives detection priorities. Strategically, to inform where the security programme invests. The common failure is subscribing to feeds and pouring indicators into the SIEM, which generates volume and little insight. Intelligence is only useful when it changes a decision, and a feed nobody acts on is a subscription rather than a capability. In the UK, NCSC advisories and sector-specific sharing communities are usually more actionable than commercial feeds, and free.

What they are assessing

Whether you can distinguish intelligence from indicator feeds.

What is the difference between an IOC and a TTP?entrytechnicalCommonly asked

An indicator of compromise is a specific artefact left behind by an attack — a file hash, an IP address, a domain, a registry key. A tactic, technique or procedure describes behaviour: not the address they used, but the fact that they establish persistence through scheduled tasks, or move laterally using valid credentials over SMB. The distinction matters because of how easily each can be changed. An attacker can swap an IP, recompile a binary to change its hash, or register a new domain in minutes — so indicators are trivially altered and expire fast. Changing their actual technique — how they persist, escalate, move — is expensive and often constrained by their tooling and skill. This is why detections built on indicators are cheap to write but brittle: you block yesterday’s hash and the same actor walks straight back in tomorrow with a new one. Detections built on behaviour are harder to write but far more durable, because they catch the thing the attacker cannot easily stop doing. The framing an interviewer wants is David Bianco’s “Pyramid of Pain”: the higher up the pyramid your detection sits — from hashes at the bottom to TTPs at the top — the more it costs the adversary to evade you. Most SOCs are bottom-heavy: thousands of indicator-based rules generating noise, and too few behavioural detections that would actually survive contact with a competent attacker.

What they are assessing

Conceptual clarity that underpins detection strategy.

How do you stay effective during a long shift working a repetitive queue?entrycompetency

By accepting that consistency matters more than intensity — the job is a marathon of small correct decisions, not a sprint of heroics. Practically: work the queue in a deliberate order rather than cherry-picking the interesting alerts and letting the dull ones rot, because the boring one is statistically as likely to be the real intrusion. Take real breaks, because sustained attention degrades measurably over a shift and a fatigued analyst miscategorises alerts they would have caught fresh — pushing through is a false economy paid for in missed detections. Lean on checklists and playbooks for routine triage so that when concentration dips, accuracy does not go with it. And know when to hand over rather than grind out the last hour of a night shift on willpower; a clean handover beats a tired solo finish. The honest core of a strong answer is a mindset point: the repetitive queue is not the boring part of the job you endure until the real work arrives — it *is* the work, and it is where most real attacks are actually caught. The alert that mattered almost always looked exactly like the two hundred benign ones around it, which is precisely why staying sharp through the monotony is the skill, not a chore around the skill.

What they are assessing

Realism about shift work, which is where new analysts most often struggle.

How would you handle a handover mid-incident?entrycompetency

A handover mid-incident is where continuity is most often lost and where a good analyst proves their discipline. The core is a clear, written state-of-play the incoming analyst can act on immediately without re-deriving everything: what is known, what has been done, what is still open, and what the immediate next actions are. Concretely — a timeline of confirmed facts with timestamps, the current containment status (what is isolated, what is not), any actions in flight that must not be duplicated or interrupted, the working hypothesis, and the explicit next step so nothing stalls in the gap. Verbal handover alone is where things fall through: memory is lossy and the outgoing analyst is tired, so it goes in the incident record, not just the conversation. The incoming analyst should read it back and ask questions before the other person leaves, because the moment to catch a misunderstanding is while both people are still present. Two things people forget: hand over the *reasoning*, not just the facts, so the next analyst does not silently re-tread a path you already ruled out; and flag anything you are unsure about explicitly, because an uncertain finding presented as settled is how incidents go sideways across a shift boundary.

What they are assessing

Whether you understand handover as a genuine risk point in shift work.

What makes incident documentation good?entrycompetency

Good incident documentation lets someone who was not there reconstruct what happened, what you did, and why — during the incident for a colleague picking it up, and long after for the lessons-learned review, an audit, or a regulator. Three properties matter. First, timeline with timestamps: what was observed and what action was taken, in order, so the attacker’s activity and the responder’s activity can be told apart later — an undocumented containment action can look exactly like adversary movement in the logs. Second, facts separated from interpretation: record what you actually saw versus what you inferred, because a hypothesis written as a fact sends the next reader down a false trail. Third, decisions and their reasons, not just actions — “isolated host X at 14:02 because it was beaconing to a known-bad domain” is usable; “isolated host X” is not. Write it as you go rather than reconstructing afterwards, because memory reorders events and softens uncertainty. The test of good documentation is blunt and worth stating: could a competent colleague pick up your incident record cold and continue without needing you in the room? If not, it is a personal aide-mémoire, not incident documentation — and the difference matters most at exactly the moments you are least available.

What they are assessing

Documentation discipline, which separates analysts who scale from those who do not.

A senior stakeholder demands an update mid-incident and you do not yet know the answer. What do you say?midscenario

First, recognise the tension honestly: the stakeholder needs assurance and information, and you need to keep working the incident — and satisfying the first by inventing detail for the second is the trap. So give them what is genuinely known without speculating beyond it. A strong holding update has a shape: what is confirmed, what is being done right now, what is not yet known, and when the next update will come. That last part does the heavy lifting — committing to a next update time (“I’ll have more at half past”) is what actually calms a demanding stakeholder, because most of the pressure is fear of being left in the dark, not impatience for a resolution you cannot yet give. Resist the pull to over-promise or to guess at cause, impact, or timeline to fill the silence; a confident wrong answer under pressure is far more damaging than an honest “we don’t know yet, here’s how we’ll find out,” because people will anchor on your number and you will spend the rest of the incident walking it back. If the interruptions themselves are impeding the response, it is legitimate and mature to say so and to route updates through an incident lead or a set cadence, protecting the people doing the technical work — knowing when to interpose that structure is a sign of seniority, not evasion.

What they are assessing

Composure and communication discipline under pressure.

You suspect an incident but the evidence is ambiguous. Do you escalate?entryscenarioCommonly asked

This tests calibration — the willingness to act under uncertainty without either crying wolf or freezing. The honest answer is that you rarely get clean evidence at the start of anything real; ambiguity is the normal condition, not a reason to wait. So the move is proportionate action, not a binary declare/ignore: investigate further to reduce the uncertainty, and take low-cost protective steps that are cheap to reverse if you are wrong — capturing volatile evidence, watching the host more closely, quietly checking related systems — while you gather more. What you should *not* do is either raise a full incident on a hunch and burn the team’s credibility, or sit on a real signal because you could not prove it yet. If genuinely unsure whether it clears the bar, escalate the *question*, not a conclusion: “here’s what I’m seeing, here’s why I’m unsure, can we look together” is a strong move, not a weak one, and a healthy SOC rewards it rather than punishing the false alarms that inevitably come with it. The failure mode the interviewer is probing for is the analyst who does nothing because the evidence was not conclusive — because in hindsight the ambiguous early signal is very often the only warning the incident ever gave.

What they are assessing

Escalation judgement, which is the single most-tested instinct in junior SOC interviews.

What would you do if you escalated something and felt it was ignored?midcompetency

This is really a question about moral courage and how you handle being overruled, and the strong answer holds two things at once: you raised it, and you respect that the decision was not yours to make alone. First, make sure the concern was actually *heard* rather than merely mentioned — put it in writing, state the specific risk and why it matters plainly, and confirm the decision-maker understood what you were flagging, because “I felt it wasn’t taken seriously” is sometimes a communication gap rather than a dismissal. If, having been clearly understood, they still judge it a lower priority, that can be legitimate: they may hold context you don’t — competing risks, business constraints, work already in train. So you document that you raised it and what was decided, and you commit to the outcome rather than relitigating it or quietly seething. But there is a floor: if the risk is genuinely serious — not merely deprioritised but dangerous — escalating past the person, through a risk-acceptance process or to someone more senior, is legitimate and sometimes obligatory. The maturity the interviewer is listening for is the judgement to tell those two cases apart: knowing the difference between “I was overruled by someone entitled to overrule me” and “this is serious enough that being overruled is not the end of my responsibility.”

What they are assessing

Persistence and professionalism when your judgement is overruled.

What is alert fatigue, and how would you address it?midtechnicalCommonly asked

The degradation in attention and judgement that comes from working a queue where almost everything is benign — analysts start closing alerts on pattern rather than analysis, and the real one gets closed with the rest. It is a detection engineering problem wearing a human costume: the fix is reducing the noise rather than exhorting people to concentrate. Practically, that means tuning the worst-offending rules, automating enrichment so each alert takes less effort, suppressing known-benign patterns properly, and measuring the false-positive rate as a first-class metric rather than an accepted cost.

What they are assessing

Whether you treat analyst attention as a finite resource to be engineered around.

How would you onboard a new log source properly?midtechnical

Establish the purpose first — what detections or investigations this source enables — because sources onboarded without a purpose become cost. Then confirm the data: what fields arrive, in what format, with what timestamps and timezone, and whether anything sensitive needs handling carefully. Then normalise it so it can be correlated with everything else rather than sitting in its own island. Then build or update the detections that justified it, and test them. Finally, add ingestion monitoring, so you find out when it stops. Skipping the last step is how estates end up with sources everyone assumes are working.

What they are assessing

End-to-end thinking rather than treating onboarding as a plumbing task.

What would you want to know in your first week in a new SOC?midcompetency

What the crown jewels are, because everything else is prioritised against them. What the detection coverage actually is rather than what the documentation claims. Where the playbooks are and how current they are. What the escalation path looks like out of hours, since that is when you will first need it. Who owns what outside the SOC, because containment usually requires someone else’s cooperation. And the recent incident history, which tells you more about the real threat picture and the organisation’s maturity than any briefing document will. In a regulated UK firm, also find out what is reportable to whom and how fast — the ICO, the FCA, or a sector regulator — because that shapes escalation more than any internal policy.

What they are assessing

Orientation instinct and awareness that a SOC exists inside an organisation.

Your SOC gets an alert at 3am for a critical server. The on-call engineer is not answering. What do you do?entryscenarioCommonly asked

The trap in this scenario is the assumption baked into “the owner isn’t answering” — that you are stuck until they do. You are not. The alert is critical and the server is critical, so the response cannot wait on one unreachable person; the job is to act within your authority while widening the net for the authority you lack. Practically, in parallel: begin triage on what you *can* see now — is the alert corroborated by other signals, is there active harm in progress — and escalate through the proper path rather than repeatedly ringing one silent phone. That means the on-call chain, the owner’s manager or team, the incident process — critical systems have escalation routes precisely for when the primary contact is unreachable at 3am, and using them is correct, not an overreach. If there is active, spreading harm, taking a cheap reversible protective step within your remit — heightened monitoring, or network isolation if the situation clearly warrants and your role permits — can be the right call, with the reasoning documented. What you must not do is the thing fatigue and deference tempt at 3am: log that you could not reach the owner and quietly close or park it. A strong answer shows you know that “I couldn’t contact one person” is never the end of the response for a critical alert — it is the trigger to escalate, not to stop.

What they are assessing

Whether you act within your authority under pressure rather than freezing or overreaching.

An analyst on your team has closed a batch of alerts unusually fast. How do you handle it?seniorscenario

The uncomfortable thing this question raises is that the explanation is not necessarily innocent, and a manager has to hold that possibility without leaping to it. Closing a batch of alerts unusually fast has several readings: the analyst found a genuine reason they were all benign (a known false-positive pattern, a bulk misfire); they are overwhelmed and cutting corners to clear a queue; they lack the knowledge to investigate properly and are closing what they don’t understand; or, least comfortably, they are deliberately hiding something. You do not assume the worst, but you do not ignore it either — you look. Start with the evidence rather than the person: pull a sample of the closed alerts and review whether the closures were actually justified, because that tells you which story you are in before you have a conversation. Then talk to the analyst in a way that opens rather than accuses — “I noticed these went through quickly, walk me through your thinking” — which surfaces an innocent explanation if there is one and applies pressure if there isn’t. If the closures were wrong, the response depends on cause: coaching and workload relief for the overwhelmed or under-skilled, something far more serious for deliberate concealment, which edges into insider-risk territory. The judgement the interviewer is testing is proportionality — taking it seriously enough to verify, without pre-judging a colleague on a pattern that often has a mundane cause.

What they are assessing

Judgement about people, and whether you see quality problems as system problems first.

You find evidence that an incident began three months ago. How does that change your response?seniorscenario

Substantially. A three-month dwell time means the attacker has had time to establish persistence in multiple places, harvest credentials, and understand the environment better than a quick sweep will reveal — so containment has to assume more than the host in front of you. It also changes the scoping question from what happened to what has been happening, which means log retention becomes the limiting factor and you may not have the data. And it changes the reporting picture: three months of potential data access is a materially different conversation with legal and the regulator than three hours.

What they are assessing

Whether you understand how dwell time changes both the technical and the organisational response.

A user calls to say they think they did something stupid. How do you handle the call?entryscenarioCommonly asked

The single most important thing here happens in the first ten seconds of the call: your tone. Someone who thinks they have done something stupid and rings the SOC anyway is doing exactly the behaviour you desperately want to encourage across the whole organisation — self-reporting early — and if that call is met with judgement or alarm, they and everyone they talk to afterwards will hesitate next time, which is how a five-minute problem becomes a five-week one. So you reassure first and mean it: thank them for calling, make it clear reporting was the right move, and get the facts without making them feel worse. Then, practically: find out what actually happened — clicked a link, entered credentials, opened an attachment, sent data to the wrong place — because the response differs completely, and take protective action proportionate to it, such as resetting credentials if they were entered somewhere, isolating the machine if something was executed, and checking for follow-on activity. The user is a witness and an ally now, not a suspect: they can tell you exactly what they saw, which is often faster than reconstructing it from logs. The interviewer is listening for whether you treat human-reported incidents as the gift they are — the fastest detection you will ever get is a person telling you directly — rather than as a failure to be scolded.

What they are assessing

Whether you understand user reporting as a detection capability that culture can destroy.

Governance, Risk and Compliance

42
How do you actually calculate risk?entrytechnicalCommonly asked

Most practical frameworks express it as a function of likelihood and impact, scored on defined scales and often plotted on a matrix. The honest position to take in an interview is that the arithmetic is the easy part and the inputs are where the judgement lives: likelihood is usually an informed estimate rather than a measured frequency, and impact depends entirely on how you scope the consequence. A five-by-five matrix produces a number that looks objective and is not. It is still useful, because it forces consistency and comparison, but a candidate who presents it as calculation rather than structured judgement has misunderstood it.

What they are assessing

Whether you can use risk scoring while being honest about what it is.

What are the four risk treatment options?entrytechnicalCommonly asked

Treat, tolerate, transfer, terminate — sometimes taught as mitigate, accept, share, avoid. Treat means applying controls to reduce likelihood or impact. Tolerate means accepting the risk as it stands, which should be a documented decision at an appropriate level rather than a default. Transfer means shifting some consequence elsewhere, usually through insurance or contract, and the point most candidates miss is that transfer moves financial impact but rarely moves accountability — the regulator and your customers still come to you. Terminate means stopping the activity that creates the risk, which is the least used and sometimes the correct answer.

What they are assessing

Whether you know the options and understand the limits of transfer.

What is risk appetite, and how does it differ from risk tolerance?midtechnicalCommonly asked

Appetite is the amount and type of risk an organisation is willing to take in pursuit of its objectives — a strategic statement, owned by the board, that sets direction. Tolerance is the acceptable variation around that: how far a specific exposure can drift outside appetite before it must be escalated or acted upon. The metaphor that makes it stick: appetite is the destination you’re driving to; tolerance is the width of the road — how far you can wander from the centre line before you’re in trouble. A worked example: a bank’s appetite might be “we will not accept risks that could materially disrupt customer payments,” while the tolerance translates that into something operable — “no more than X minutes of payment downtime per quarter before board escalation.” That translation is the whole point, and it’s where most of the real GRC work lives. In practice, a great many organisations have appetite statements written at a level so abstract (“we have a low appetite for cyber risk”) that they cannot be used to make a single decision, and the valuable work is turning them into thresholds that actually mean something operationally — which systems, which data, which outage duration, which loss figure. A strong interview answer names that gap: an appetite statement nobody can apply to a real decision is decoration, and the skill is building the ladder from board-level appetite down to the operational thresholds an analyst can act on at 3am.

What they are assessing

Precision about two terms that are constantly conflated, and awareness that appetite statements are often useless in practice.

What makes a risk acceptance decision valid?midtechnicalCommonly asked

Four things, and an acceptance that fails any one of them is not a governance decision. First, authority: the decision is made by someone with the delegated authority to accept a risk of that size, which requires a defined delegation structure rather than whoever happened to be in the room when it came up — a team lead cannot validly accept a risk that could sink the company. Second, informed consent: the person accepting genuinely understands what they are accepting, expressed in business terms (“we could lose customer data for two days during month-end”) rather than technical ones (“the CVE is unpatched”), because you cannot validly accept a risk you do not understand. Third, documentation: the decision, the reasoning, and the residual exposure are written down — partly so it is auditable, partly because the act of writing it forces clarity about what is actually being accepted. Fourth, an expiry: a review date, because circumstances change and an acceptance with no end date silently becomes permanent, describing a world that may no longer exist. The reason this matters beyond box-ticking is accumulation — organisations accrue dozens of stale, forgotten acceptances that collectively describe a risk posture nobody ever actually chose. A valid acceptance is a deliberate, authorised, time-boxed decision; an invalid one is, as the saying goes, just a note that somebody once shrugged.

What they are assessing

Whether you understand acceptance as governance machinery rather than paperwork.

What is a risk register actually for?entrytechnicalCommonly asked

To make risk visible, comparable and owned — three jobs, each of which fails in a recognisable way when the register is treated as an artefact rather than an instrument. Visible, so decisions are made deliberately rather than by drift: a risk nobody has written down is a risk nobody is deciding about, it is just happening. Comparable, so that limited money and attention flow to the largest exposures rather than the loudest voices or the newest fear — a register’s scoring exists precisely to let you rank a boring high-impact risk above an exciting low-impact one. Owned, so that every entry has a named person accountable for doing something, because a risk owned by “IT” or “the business” is owned by no one. The failure mode is treating the register as the objective in itself: organisations produce beautiful, comprehensive, colour-coded registers that change nothing, because the entries have no real owners, no dates, and no consequence for sitting untouched quarter after quarter. The blunt test an interviewer is listening for: has this register actually caused a decision — money moved, a project delayed, a control funded — in the last year? If not, it is documentation, not governance, and the sophistication of its heat-map is beside the point.

What they are assessing

Whether you see the register as an instrument or an artefact.

How would you get business owners to engage with risk they see as IT’s problem?midcompetencyCommonly asked

By translating out of security language and into theirs, because ownership follows comprehension — people will not own a risk they experience as somebody else’s technical problem. A risk described as “unpatched critical vulnerabilities in the CRM” unmistakably belongs to IT; the *same* risk described as “we could lose access to customer records for several days during month-end close” unmistakably belongs to the person who owns month-end. Nothing changed but the framing, and the framing is what moves the risk from the security function to the business — which is the core GRC skill and the hardest to fake. Then make the ownership concrete and light: one named person, one specific decision to make, and a short conversation rather than a form to complete, because effort is the enemy of engagement. And show the consequence either way — what happens if we treat this, what happens if we accept it — so the decision feels real rather than bureaucratic. Engagement usually fails for three compounding reasons: the ask is abstract (“engage with risk”), it is effortful (a spreadsheet to fill in), and it is framed in a vocabulary the owner does not use. Fix those three and most reluctant owners engage, because you have turned an IT chore into a business decision that is visibly theirs.

What they are assessing

Whether you can move risk from the security function to the business, which is the core GRC skill.

What is the difference between a risk and an issue?entrytechnical

A risk is something that might happen; an issue is something that already has. The distinction sounds pedantic but it reveals immediately whether someone has worked in a real governance process, because the two route completely differently. Risks go into the register to be assessed, scored and given a treatment decision — accept, treat, transfer, avoid. Issues go straight into remediation with an owner and a due date, because there is nothing left to assess about whether they will happen: they have. Conflating them produces two specific, common failures. First, registers clogged with things that are already true — “we have no MFA on the VPN” is not a risk to be scored on likelihood, it is an issue to be fixed, and scoring its probability is faintly absurd. Second, and more dangerous, live problems being managed at the leisurely cadence of a quarterly risk review, when they need an owner and a deadline this week. The practical tell of a mature process is that the moment something crosses from “might” to “has,” it changes track: stop scoring its likelihood and start fixing it, with a name and a date attached.

What they are assessing

Basic precision that reveals whether you have worked in a real governance process.

How would you build a risk assessment methodology from scratch?seniortechnical

Start with what decisions it needs to support, because a methodology that produces scores nobody acts on is overhead. Then define the scales — likelihood and impact — in language specific to the organisation, with impact expressed across the dimensions that matter to it: financial, regulatory, customer, operational. Define what each score level means concretely so two assessors reach similar answers. Set the treatment thresholds and the delegation of acceptance. Then pilot it on a handful of real risks and adjust, because methodologies written in the abstract always produce surprises on contact. Simplicity beats sophistication: a method people use badly beats one they avoid.

What they are assessing

Whether you can design governance rather than only operate it.

What does the ICO actually expect after a personal data breach?midtechnicalCommonly asked

Notification to the ICO without undue delay and, where feasible, within 72 hours of *becoming aware* — where “aware” means a reasonable degree of certainty that a security incident occurred and personal data was affected, not the completion of your investigation. The 72 hours is calendar time, including weekends, and the clock starts at awareness, not at the breach itself, which candidates routinely get wrong. Crucially, not every breach is reportable: the threshold is whether it is *likely to result in a risk to the rights and freedoms of individuals*. A breach of properly encrypted data whose keys were not compromised may well fall below that bar; a breach of unencrypted health or financial data almost certainly clears it. If you assess a breach as not reportable, document that reasoning, because the ICO can ask you to justify it and “we decided it was fine” is not an answer. Two further points a strong candidate raises unprompted. First, there is a separate, higher threshold for telling the affected *individuals* — “high risk” to their rights and freedoms, not merely “risk” — so plenty of breaches are reportable to the ICO but need not be communicated to individuals. Second, you can notify in phases: if you don’t have all the facts within 72 hours, submit an initial notification with reasons for any delay and supplement it later. (For interview currency: these thresholds were not changed by the Data (Use and Access) Act 2025.)

What they are assessing

Whether you know the actual trigger and the actual threshold, which candidates routinely get wrong.

What is a Data Protection Impact Assessment, and when is one required?midtechnicalCommonly asked

A structured assessment of how a processing activity affects individuals’ privacy rights, and what will be done to reduce that impact. It is required where processing is likely to result in a high risk to individuals — typically large-scale processing of special category data, systematic monitoring of publicly accessible areas, automated decision-making with significant effects, or use of new technologies in ways individuals would not expect. The practical point for a security professional is that you are usually a contributor rather than the owner: the DPO or privacy team runs it, and you supply the security control picture.

What they are assessing

Whether you know your role in the process rather than claiming to own it.

What is the role of a Data Protection Officer, and when must one be appointed?midtechnical

To advise on data protection obligations, monitor compliance, act as the contact point for the ICO and for individuals, and do so with genuine independence — the role cannot be instructed on how to reach its conclusions and cannot be penalised for them. Appointment is mandatory where the organisation is a public authority, or where core activities involve large-scale systematic monitoring or large-scale processing of special category data. Many organisations appoint one voluntarily. The conflict-of-interest point is examinable: the DPO generally should not be the person deciding the purposes and means of processing.

What they are assessing

Whether you understand the independence requirement, which is the part most often misunderstood.

What are the NIS Regulations, and who do they apply to?midtechnical

UK legislation imposing security and incident-reporting duties on operators of essential services — energy, transport, health, water, digital infrastructure — and on relevant digital service providers. Duties include taking appropriate and proportionate technical and organisational measures, and reporting significant incidents to the designated competent authority for the sector, which differs by sector rather than being a single regulator. For security professionals the practical relevance is that NIS-regulated organisations are typically assessed against the NCSC’s Cyber Assessment Framework, so the two come as a pair.

What they are assessing

UK regulatory literacy beyond GDPR, which distinguishes candidates who have worked in regulated sectors.

What is the Cyber Assessment Framework, and how is it used?midtechnical

The NCSC’s framework for assessing cyber resilience, structured around four objectives — managing security risk, protecting against cyber attack, detecting cyber security events, and minimising the impact of incidents — broken into principles and contributing outcomes. Rather than a control checklist, it is outcome-based: you evidence that an outcome is achieved, not that a specific control exists. It is the assessment basis for NIS-regulated organisations and is widely adopted across UK government and critical national infrastructure, which makes naming it a strong signal in public-sector interviews.

What they are assessing

Whether your framework knowledge includes the British ones, not only NIST and ISO.

What is the difference between the FCA and the PRA?midtechnicalCommonly asked

Both regulate UK financial services and their remits differ. The FCA is the conduct regulator, concerned with how firms behave toward consumers and markets, and it regulates a very large number of firms. The PRA, part of the Bank of England, is the prudential regulator, concerned with the safety and soundness of the firms whose failure would threaten financial stability — banks, building societies, insurers and major investment firms. Larger firms are dual-regulated. For security professionals the practical point is that both have expectations on operational resilience and cyber, and a dual-regulated firm answers to both.

What they are assessing

Basic UK financial-services literacy, expected in the sector’s largest security employer.

What is DORA, and does it affect UK firms?seniortechnical

The EU’s Digital Operational Resilience Act, in application since 17 January 2025, which sets binding requirements on financial entities across ICT risk management, incident reporting, digital operational resilience testing, and third-party ICT risk — including, notably, direct EU oversight of the critical ICT providers the sector depends on, such as major cloud platforms. Does it affect UK firms? Directly, no — it is EU law and the UK is outside it post-Brexit. But the honest, senior-level answer is that “doesn’t apply directly” is not the same as “doesn’t matter,” and the ways it reaches UK firms are exactly what an interviewer is probing for: a UK group with EU-authorised entities, EU customers, or EU-regulated group members will be pulled into DORA’s scope through those, and many UK groups have chosen to align their whole estate rather than run two parallel resilience frameworks. The genuinely useful thing to demonstrate is that you understand the relationship rather than just the acronym: DORA and the UK’s own operational resilience regime (the FCA/PRA rules on important business services and impact tolerances) are addressing the *same* underlying concern — that the financial system has become critically dependent on a handful of technology providers — with different mechanics and timelines. Knowing that lets you talk about it as a coherent regulatory direction of travel rather than a foreign rule you can safely ignore.

What they are assessing

Whether you track regulation beyond the UK border, which matters in any firm with European operations.

What is the Computer Misuse Act, and why should a security professional care?midtechnical

The UK legislation criminalising unauthorised access to computer material, unauthorised access with intent to commit further offences, and unauthorised acts impairing operation. Security professionals should care for two reasons. First, it is the legal basis on which attackers are prosecuted. Second and more practically, it is what makes authorisation the dividing line for your own work: testing without documented permission is potentially a criminal offence regardless of intent, which is why scope documents and rules of engagement matter so much. The Act has long been criticised for lacking a clear defence for legitimate security research.

What they are assessing

Legal awareness, and understanding that authorisation is what separates the profession from the offence.

What frameworks would you use, and why?midtechnicalCommonly asked

Different frameworks answer different questions, so the honest answer names a few with their purposes rather than listing everything. NIST CSF for structuring and communicating a programme, because its five functions give non-specialists something to hold. ISO 27001 where a certifiable management system is needed. NCSC guidance and the CAF in UK public sector and critical services. CIS Controls where a prioritised technical starting point is wanted. What matters more than the list is using them as references rather than scripture: a framework describes what good looks like in general, not how to get there in your specific estate.

What they are assessing

Whether you use frameworks or collect them — and whether you know the British ones.

What is the difference between ISO 27001 and ISO 27002?midtechnicalCommonly asked

ISO 27001 is the certifiable standard: it specifies the requirements for an information security management system, and it is what an organisation is audited against. ISO 27002 is guidance — it provides implementation advice for the controls referenced in 27001’s Annex A, and you cannot be certified against it. In practice 27001 tells you that you must select and justify controls based on risk; 27002 helps you work out what each control actually means. Candidates who describe 27002 as a certification, or who cannot separate the management system from the control set, reveal that they have read about the standard rather than worked with it.

What they are assessing

Precision about a pair that candidates routinely conflate.

What is a Statement of Applicability?midtechnical

The document recording which Annex A controls apply to the organisation, which do not, and the justification for each decision — including justification for exclusions. It is a mandatory part of ISO 27001 certification and is one of the first things an auditor examines, because it reveals whether control selection was risk-driven or copied. The common weakness is a Statement of Applicability that includes everything with no reasoning, which signals that nobody performed the risk assessment it is supposed to reflect. Exclusions are entirely legitimate if justified; unjustified inclusions are the actual red flag.

What they are assessing

Whether you have worked with an ISMS rather than only read about one.

What is SOC 2, and how does it differ from ISO 27001?midtechnical

SOC 2 is an American attestation report produced by an accredited auditor, assessing controls against the Trust Services Criteria — security, and optionally availability, processing integrity, confidentiality and privacy. ISO 27001 is an international certification against a management system standard. Three practical differences: SOC 2 produces a detailed report you share with customers under NDA, while ISO produces a certificate; SOC 2 Type II covers a period of operation whereas Type I is a point in time; and SOC 2 is what North American customers usually ask for, while ISO 27001 is what European and UK customers expect.

What they are assessing

Whether you can advise on which assurance a customer actually wants, which is a real commercial question.

What is PCI DSS, and when does it apply?midtechnical

The Payment Card Industry Data Security Standard applies to any organisation that stores, processes or transmits cardholder data, and is enforced contractually by the card brands and acquiring banks rather than by statute. Requirements are prescriptive rather than risk-based, which distinguishes it from ISO 27001. The single most valuable concept for an interview is scope reduction: the standard applies to the cardholder data environment, so architectural choices that keep card data out of your systems — tokenisation, redirect or iframe payment pages, outsourcing to a compliant provider — shrink both the compliance burden and the risk.

What they are assessing

Whether you understand scope reduction, which is the practical heart of PCI work.

How would you prepare an organisation for its first ISO 27001 certification?seniortechnical

Scope first, and deliberately — an over-broad scope makes the first certification far harder than it needs to be, and scope can widen later. Then the management system itself: leadership commitment, policy, roles, and above all a working risk assessment, because everything downstream derives from it. Then the Statement of Applicability, then closing the control gaps it reveals. Then the parts organisations forget until late: internal audit, management review, and evidence of continual improvement, all of which need to have actually happened before the certification audit rather than being described as intentions.

What they are assessing

Whether you know the sequence and the commonly missed requirements.

How would you write a security policy people actually follow?midtechnicalCommonly asked

Make it short, specific, and about what people must do rather than what the organisation aspires to. Long policies are unread policies, and unread policies are unfollowed ones. Write for the audience: an acceptable use policy is read by everyone in the organisation and should be plain English, while a technical standard can assume expertise. State the reason briefly, because compliance improves when people understand why. Make the compliant path the easy one, because a policy that requires heroic effort will be routed around. And review on a real cadence, since a policy referencing systems retired years ago teaches people the whole set is fiction.

What they are assessing

Whether you write policy to change behaviour or to satisfy an auditor.

How would you measure whether a policy is working?midtechnical

Not by attestation rates — “98% of staff have read the policy” measures whether people clicked a button, not whether the policy changes anything. The meaningful measures are behavioural and control-based. How often is the policy actually breached in practice? How many exceptions have been requested and granted? Are the technical controls that enforce it genuinely operating, or has enforcement quietly drifted? Exception volume is the single most informative signal and the one people overlook: a policy generating a constant stream of exception requests is usually *wrong* rather than widely disobeyed — it is asking for something unreasonable in the real operating environment — and the correct response is to fix the policy, not to keep granting exceptions or to crack down on the people requesting them. Alongside that, run spot checks against reality: does what the policy *requires* match what the systems actually *do*? A password policy that mandates settings the identity platform doesn’t enforce is a policy that exists on paper and nowhere else. The mindset an interviewer is listening for is that a policy is an intervention with a measurable effect on behaviour and risk, not a document to be signed and filed — and if you can’t point to how you’d know whether it’s working, you can’t claim it is.

What they are assessing

Whether you treat policy as something with measurable effect rather than a document to be signed.

How would you handle a policy exception request?midtechnicalCommonly asked

Treat it as a small risk acceptance, because that is exactly what it is — and framing it that way immediately imports the discipline that makes exceptions safe. First, understand what is actually being asked and *why the policy cannot be met*, because sometimes the honest finding is that the policy is unreasonable for this case and the right fix is to change the policy rather than grant a one-off. Then assess the residual risk with any compensating controls the requester can put in place — an exception with a mitigating control is a very different proposition from a bare one. Route the decision to whoever holds the authority to accept a risk of that size, the same delegation logic as any risk acceptance. And then the part that matters most and is skipped most often: put an expiry on it. Exceptions without end dates are how policy erodes — they accumulate silently until the policy describes a world that no longer exists, and every live exception is a documented hole that, past its useful life, nobody is watching or revisiting. The accumulation problem is the real risk here, not the individual grant: ten reasonable exceptions, each sensible on its own, can add up to a control that has effectively been switched off, and only a mandatory review date catches that.

What they are assessing

Whether you connect exceptions to risk acceptance and understand the accumulation problem.

What is the three lines model?midtechnicalCommonly asked

A governance structure that separates responsibility for risk into three distinct roles so that the people doing the work, the people setting the rules, and the people checking are not the same people. The first line is the business — the teams that own and manage their risks day to day as a direct part of doing their jobs. The second line is the oversight functions — risk, compliance, and often security governance — which set the framework, provide expertise, and challenge the first line, but do not own the risks themselves. The third line is internal audit, which provides *independent* assurance to the board that the first two lines are actually working, and which reports to the audit committee rather than to management precisely so it can say uncomfortable things without being overruled by the executives it is assessing. The whole point is that independence: assurance is only worth anything if the assurer can be honest, and an audit function that reports to the people it audits cannot be. Two nuances a stronger answer adds: the model is a way of thinking about accountability, not a rigid org chart, and smaller organisations blend the lines out of necessity — but even then, someone should be able to articulate who owns, who oversees, and who independently checks, because when all three collapse into one team, no one is really assuring anything.

What they are assessing

Whether you understand where a GRC function sits and why independence matters.

How would you prepare for an external audit?midtechnicalCommonly asked

Know the scope and the standard you are being assessed against, and map the evidence to it before the auditor arrives rather than during. Run an internal audit first, because finding your own gaps is far cheaper than having them found for you. Make sure evidence exists in the form auditors accept — records that something happened on a date, not assurances that it usually does. Brief the people who will be interviewed so they answer honestly and concisely rather than volunteering tangents. And be straight about known gaps: auditors respond far better to a known issue with a remediation plan than to a discovered one.

What they are assessing

Practical audit experience, particularly the instinct to self-identify gaps.

An auditor raises a finding you believe is wrong. How do you handle it?midcompetencyCommonly asked

Separate the fact from the conclusion. Often the observation is correct and the rating or the implication is what you disagree with, and that is a much more productive conversation than disputing the whole finding. Provide evidence rather than argument — if the control was operating, show the records. If you still disagree after that, most audit processes allow a documented management response recording your position alongside the finding, which is the right mechanism. What damages you is emotional pushback, or accepting a wrong finding to avoid friction and then owning a remediation action for a problem that does not exist.

What they are assessing

Professional handling of disagreement with an assurance function.

How would you assess a critical supplier?midtechnicalCommonly asked

Proportionately, and starting from exposure rather than from a questionnaire. What will they access, hold or connect to, and what happens to us if they fail or are breached? That determines depth. For a genuinely critical supplier: their security certifications and what those actually cover, their incident history and how they handled it, their own supply chain, the contractual terms on notification and audit rights, and the exit plan. The question I would want answered above all is what happens to our data and our service if they are compromised on a Friday night — because that is the scenario the assessment exists for.

What they are assessing

Proportionality and outcome focus rather than questionnaire ritual.

What is fourth-party risk, and why does it matter?seniortechnical

The risk arising from your suppliers’ suppliers — the dependencies you have no contract with and often no visibility of. It matters because concentration hides there: several of your critical suppliers may depend on the same cloud region, the same payments processor, or the same identity provider, so an outage or compromise that looks unrelated on your supplier list is actually a single point of failure. The practical approach is to ask critical suppliers to disclose their own critical dependencies, and to look for concentration across your portfolio rather than assessing each supplier in isolation.

What they are assessing

Whether your supply-chain thinking extends past the first tier.

How would you handle a supplier who refuses to complete your security assessment?midscenario

Find out *why* before you escalate, because the refusal is usually informative rather than obstructive. Large suppliers with strong market positions frequently decline bespoke security questionnaires as a matter of standing policy — they cannot complete a different 200-question spreadsheet for every customer — and instead offer standardised assurance: ISO 27001 certification, a SOC 2 Type II report, penetration test summaries, published security documentation. That standardised evidence is often *better* than a self-completed questionnaire, not worse, because it has been independently assessed rather than self-attested by the very party you are trying to assess. So the first move is to check whether what they offer actually covers your specific concerns. If it does, accept it — insisting on your form for its own sake is process rigidity, not diligence. If it genuinely does not — they can’t evidence something that matters for your use case — then the decision becomes an explicit business one: state the residual risk of proceeding without that assurance plainly, and route it to whoever owns that risk to accept or reject. What you must not do is either wave the supplier through quietly because they’re big and pushing back is awkward, or block them unilaterally on principle. The judgement being tested is whether you can tell the difference between assurance you actually need and a questionnaire you’re attached to.

What they are assessing

Pragmatism and correct escalation rather than process rigidity.

What is the difference between RTO and RPO?entrytechnicalCommonly asked

Recovery Time Objective is how long a service can be down before the consequences become unacceptable — a target for how fast you must restore. Recovery Point Objective is how much data you can afford to lose, expressed as time — a target for how recent your last usable recovery point must be, which in practice sets your backup or replication frequency. A concrete way to hold them apart: if your RPO is one hour, you need a recovery point no more than an hour old, so you must be backing up or replicating at least hourly; if your RTO is four hours, you must be able to get the service running again within four hours of it failing. They drive genuinely different investments — RTO drives failover capability, warm standby, and rehearsed recovery runbooks; RPO drives replication technology and backup cadence. The common interview error is defining both correctly and then stopping, without connecting them to money. Both tighten *expensively* and non-linearly: an RPO of near-zero implies continuous replication, an RTO of near-zero implies hot standby infrastructure, and each order-of-magnitude improvement can multiply cost. The genuinely valuable GRC contribution is helping the business decide what it is actually willing to pay for each — which usually means discovering that the “we can’t lose any data and can’t have any downtime” demand evaporates once someone sees the invoice for delivering it.

What they are assessing

Whether you connect continuity targets to the cost decisions they imply.

What is a business impact analysis?midtechnical

The exercise that determines which activities matter most and how quickly they must be restored, by assessing the consequences of their disruption over time — financial, regulatory, customer, reputational. It is the input to continuity planning: without it, recovery priorities are set by whoever argues loudest. The output is a ranked picture of critical activities with their tolerable downtime and dependencies. The most common weakness is running it with IT rather than with the business, which produces a list of systems rather than a list of business activities — and the two are not the same, because a critical activity often depends on several unremarkable systems.

What they are assessing

Whether you understand BIA as a business exercise rather than a technical inventory.

How would you test a business continuity plan?midtechnicalCommonly asked

In escalating levels, because jumping straight to a full test wastes effort on problems a cheaper exercise would have found. A walkthrough confirms people know the plan exists and their role in it. A tabletop exercise tests decision-making against a scenario. A functional test exercises specific capabilities, such as actually restoring from backup. A full simulation tests the whole thing under realistic conditions. The critical discipline at every level is that findings are captured and fixed — an exercise that produces a warm feeling and no actions has tested nothing. And test the assumptions people are most confident about, because that is where the failures hide.

What they are assessing

Whether you know the testing hierarchy and treat exercises as finding-generators.

What is operational resilience, in outline?midtechnicalCommonly asked

The ability to continue delivering services through disruption, whatever the cause. It differs from traditional continuity planning in its starting point: rather than asking how we recover systems, it asks which services must keep running for customers and the market, then works inward to what they depend on. In UK financial services it is a regulatory regime with specific mechanics — identify important business services, set impact tolerances, map the people, processes, technology, facilities and third parties each depends on, and test against severe but plausible scenarios. The mindset shift is assuming disruption will happen rather than trying to prevent all of it.

What they are assessing

Whether you understand the outside-in framing that distinguishes resilience from continuity.

How would you explain a technical risk to a board?midcompetencyCommonly asked

Lead with the business consequence, not the mechanism — this is the single most valued capability in senior GRC roles and the one most technical people get wrong. A board does not need to understand the vulnerability, the exploit chain, or the CVSS score; they need to understand what could happen, how likely it is, what it would plausibly cost, and — critically — what decision is being asked of *them*. Give them a decision, not a briefing: “here is the exposure, here are two or three options with their costs and residual risk, here is my recommendation.” A board that receives information without a recommendation will either defer (and nothing happens) or improvise (and something worse happens). Use their vocabulary throughout — financial impact, regulatory exposure, customer harm, operational disruption, reputational damage — because those are the axes on which they already make every other decision, and translating cyber risk onto them is precisely the translation the role exists to perform. And be ruthlessly brief: the discipline of compressing it to two minutes and one slide forces the clarity that a twenty-minute technical walkthrough uses complexity to avoid. The senior-probe version of this: if you cannot state the risk, the options, and your recommendation in the time it takes the lift to reach the top floor, you have not yet understood it well enough to bring it to a board.

What they are assessing

Translation skill, which is the single most valued capability in senior GRC roles.

How would you present security metrics to an executive audience?seniorcompetencyCommonly asked

Choose metrics that answer questions an executive actually has: are we getting better, where is our biggest exposure, and is the investment working. That rules out most of what security teams instinctively report — attacks blocked, emails filtered, alerts closed — which measure activity rather than outcome. Show trend rather than snapshot, because direction matters more than a single number. Be honest about the bad ones, because a dashboard that is always green trains the audience to stop reading it. And attach every metric to a decision or an action, otherwise you are presenting weather rather than a proposal.

What they are assessing

Metric literacy and resistance to vanity reporting.

You inherit a risk register with 400 open items. Where do you start?midscenarioCommonly asked

Not by remediating, because a 400-item register is usually a data quality problem rather than 400 real risks. First, deduplicate and consolidate — the same underlying issue is typically recorded many times from different assessments. Then triage for validity: how many are already resolved, obsolete, or actually issues rather than risks. What remains is the real register, usually a fraction of the original. Then assign ownership, because unowned risks do not move, and prioritise by genuine exposure rather than by inherited rating. Then report honestly on what you found, because the number itself is a finding about how governance has been operating.

What they are assessing

Whether you resist the obvious action and diagnose before treating.

The business wants to launch in two weeks and the security assessment is not complete. What do you do?midscenarioCommonly asked

Work out what can be assessed in the time, rather than treating the choice as binary between “fully assessed” and “blocked.” The instinct to say “security isn’t done, so we can’t launch” is exactly the instinct that gets security excluded from the next launch, and the instinct to wave it through to be helpful is how real harm ships. The mature move is triage: focus the two weeks on what would *actually stop a launch* — exposure of customer data, a clear regulatory breach, anything catastrophic and irreversible — and consciously set aside the findings that can be assessed and remediated safely after go-live. Then present the position with complete honesty and no theatre: here is what we assessed and found, here is what we did *not* have time to assess, here is the residual uncertainty that creates, and here is the concrete risk of launching now. Attach a remediation plan with dates for the deferred items so “later” is a commitment rather than a hope. And route the go/no-go to whoever owns that risk — because launching with known, documented, time-boxed residual risk is a legitimate business decision that is theirs to make, whereas you *pretending* the assessment was complete is not. Blocking unilaterally and waving it through are failures of the same kind: both substitute your comfort for an honest, owned decision. The interviewer is testing whether you can compress scope under pressure without compressing your integrity.

What they are assessing

Whether you can compress an assessment without pretending it was complete.

How would you build a security awareness programme that changes behaviour?midtechnical

Target specific behaviours rather than general awareness, because "be more careful" is not actionable. Pick the behaviours that actually reduce risk in your organisation — reporting suspicious messages, not reusing credentials, challenging unknown visitors — and design for each. Make reporting easy and reward it visibly, because the reporting culture is worth more than the click rate. Segment the audience, since finance and developers face different risks. And measure behaviour rather than completion: training completion rates measure attendance, while reporting rates and repeat-clicker trends measure whether anything changed.

What they are assessing

Whether you distinguish awareness activity from behavioural outcome.

What would you do in your first ninety days in a new GRC role?midcompetencyCommonly asked

Understand before proposing. What the organisation does and what would genuinely hurt it, which is the basis for every risk judgement that follows. What regulatory obligations actually apply, since that is non-negotiable ground. What governance already exists and whether it functions or merely documents. Who the stakeholders are and what they think of the security function, because GRC runs entirely on relationships and credibility. Then find one visible, achievable improvement and deliver it, because credibility is earned by usefulness rather than by assessment. Arriving with a framework and a plan to implement it, before understanding any of the above, is the classic failure.

What they are assessing

Orientation instinct, and awareness that GRC influence is earned rather than granted.

What is the biggest weakness of compliance-driven security?seniorcompetencyCommonly asked

That it optimises for demonstrable conformance rather than for actual resilience, and the two diverge. A compliant organisation has evidenced that specified controls exist; it has not demonstrated that those controls work against a real adversary, that they still work after two years of drift, or that they cover the risks the standard did not anticipate. Compliance also creates a ceiling, because effort stops when the requirement is met. The honest position is that compliance is a useful floor — it forces baseline hygiene and gives security leverage it would not otherwise have — but treating a certificate as evidence of security is the error the whole profession keeps making.

What they are assessing

Whether you can hold compliance value and compliance limits at the same time.

Penetration Testing

43
What is CREST, and why does it matter in the UK?entrytechnicalCommonly asked

CREST is the UK’s principal accreditation body for technical security services, certifying both companies and individual practitioners against assessed standards of competence and process. It matters commercially because much of the serious UK market — government, financial services, critical national infrastructure — will only buy testing from CREST-accredited firms, so the accreditation is effectively a licence to compete for that work; a firm without it is locked out of whole categories of contract regardless of how good its people are. For an individual it structures the career path: CREST’s certifications ladder from entry practitioner exams up through team-leader level, and progressing them is how you evidence seniority to employers and clients. The reason this comes up in interviews is that it’s a fast test of whether you understand the *UK* market specifically. A candidate who talks only about OSCP and American certifications has prepared for a different country; knowing where CREST sits relative to the NCSC’s CHECK scheme (government work) and the Tiger Scheme (an alternative accreditation) signals that you understand how UK testing is actually bought and regulated. A strong answer places CREST in that ecosystem rather than describing it in isolation.

What they are assessing

UK market literacy, which distinguishes candidates who have prepared properly.

What is CHECK, and who needs it?midtechnical

A UK government scheme, run by the NCSC, that authorises individuals and companies to perform penetration testing of systems holding government information or supporting government services. Testers hold CHECK Team Member or CHECK Team Leader status — the Team Leader being the senior role that can sign off and lead engagements — and that status is awarded off the back of qualifying certifications through CREST or the Tiger Scheme rather than being a separate exam. The detail that trips people up, and the thing worth knowing, is that CHECK work generally requires security clearance, usually SC as a minimum and sometimes higher depending on the system. The practical career implication is significant and worth stating plainly: CHECK-eligible testers command a premium and can reach a segment of the market that is completely closed to everyone else, but the clearance takes months, must be sponsored by an employer, and cannot be obtained speculatively — so it is not something you arrive in the industry already holding. That sequencing matters for anyone planning a testing career: you get the technical certifications first, join a firm that can sponsor clearance, and the CHECK-gated work opens up after the vetting completes, not before.

What they are assessing

Whether you understand the clearance-gated segment of the UK testing market.

What is CBEST, and who is it for?seniortechnical

CBEST is a framework for intelligence-led penetration testing of critical financial infrastructure, commissioned jointly by the Bank of England and the FCA and aimed at the systemically important firms whose failure could threaten financial stability. It differs from ordinary penetration testing in two fundamental ways that are the whole point of the answer. First, it is threat-intelligence-led: rather than working through a generic methodology, accredited threat-intelligence providers first build a picture of the real, assessed threat actors targeting *that specific firm* and their actual techniques, and the test then emulates those. Second, its target is the whole organisation’s resilience — including whether the blue team detects and responds — not merely a list of technical vulnerabilities; a CBEST test that gets caught early is in one sense a *success* for the firm, which is the opposite of how a normal pentest is judged. Testing is delivered by accredited providers using senior-certified testers, precisely because emulating a real advanced adversary safely against live critical systems demands seniority. The senior-level context worth adding: its European counterpart is TIBER-EU, which follows the same intelligence-led, whole-organisation principles, and the two are part of a broader regulatory move toward testing detection and response rather than just prevention.

What they are assessing

Awareness of the top end of the UK testing market.

What is red teaming, and how does it differ from penetration testing?midtechnicalCommonly asked

A penetration test aims to find as many exploitable weaknesses as possible in a defined scope, within a known timeframe, usually with the defenders aware. A red team engagement aims to achieve specific objectives — reach this data, compromise this system — using whatever routes an actual adversary might, including social engineering and physical access, usually without the defenders knowing. The purpose differs accordingly: a pen test measures your vulnerabilities, a red team measures your detection and response. A red team that finds a path in and is never noticed has produced a more uncomfortable and more useful finding than a list of CVEs.

What they are assessing

Whether you understand the two as different products with different purposes.

What is a white cell, and why does it exist?seniortechnical

The small, tightly controlled group of people inside the target organisation who know a red team engagement is actually happening — often just a handful of senior individuals. It exists for control and safety, and a good answer names the three distinct jobs it does. First, authorisation: if defenders escalate what they think is a real intrusion to executives, legal, or even the police, someone in the white cell must be able to step in and confirm the activity is sanctioned before a genuine crisis spins up. Second, a kill switch: someone must be able to pause or stop the exercise instantly if it starts to threaten production systems or real customers, because a red team operating realistically can cause real damage. Third, deconfliction: if a *genuine* attacker happens to be active during the engagement window, the white cell is what lets the organisation tell the exercise and the real thing apart — otherwise the red team’s noise masks the real intrusion, or the real intrusion gets dismissed as “just the test.” The consequence of not having one is the point: without a white cell, a convincing red team can trigger a full-blown incident response, wasting money, exhausting the blue team, and damaging the very relationship the testing was meant to strengthen. It’s the control that keeps realistic testing from becoming an actual emergency.

What they are assessing

Operational maturity about how adversarial engagements are actually run safely.

What testing methodologies or standards do you work to?entrytechnicalCommonly asked

The honest answer names a few and explains what each is for. OWASP’s Testing Guide and ASVS for web applications — the Testing Guide gives the procedures, ASVS gives verifiable requirements at defined levels. The Penetration Testing Execution Standard for overall engagement structure. NIST SP 800-115 as a general technical assessment reference. For UK work, CREST’s own guidance shapes how accredited firms are expected to operate. What matters more than the list is using them as a completeness check rather than a script, because a methodology tells you what not to forget, not what this particular application is doing wrong.

What they are assessing

Whether you have a repeatable structure, and know the difference between the Top 10 and an actual methodology.

How would you scope a web application test?midtechnicalCommonly asked

Establish what the application is and does before counting anything: user roles, business functions, what data it holds, what integrations it has. Then the technical boundary — which hostnames and environments, whether APIs are included, whether the underlying infrastructure is in scope. Then access: will you get credentials for each role, because unauthenticated-only testing misses most access-control issues. Then constraints: test or production, permitted techniques, timing windows, and whether denial-of-service is excluded, which it almost always is. Scoping by page count or by hostname alone produces the classic failure — a quote that does not survive contact with the actual application.

What they are assessing

Whether you can scope commercially and technically rather than guessing.

What is the difference between black box, grey box and white box testing?entrytechnicalCommonly asked

Black box means no prior knowledge — you start where an external attacker would. Grey box means partial knowledge, typically credentials and some documentation. White box means full access to source, architecture and configuration. The commercial reality is that grey box usually delivers the best value: black box burns paid days on reconnaissance an attacker would happily spend weeks on unpaid, and finds less. Clients often request black box because it feels more realistic; the honest advice is that unless you are specifically testing detection, you get more security per pound from giving the tester a head start.

What they are assessing

Whether you can advise a client on test type rather than just define the terms.

How would you test for broken access control?midtechnicalCommonly asked

Systematically, with at least two accounts. Map the roles and the objects each role should reach, then attempt every crossing: horizontal, where you use account A to reach account B’s data by manipulating identifiers, and vertical, where a standard user invokes an administrative function directly. Test the API as well as the interface, because access control is frequently enforced only in the front end. Check that functions hidden from the UI are actually refused when called. The reason this class dominates real findings is that it cannot be scanned for — a tool cannot know that invoice 1004 should belong to someone else.

What they are assessing

Method rather than opportunism, and awareness of why scanners miss this class.

How would you test an API?midtechnicalCommonly asked

Start from the specification if one exists — a Swagger or OpenAPI document is the fastest route to complete coverage, and its absence is itself worth noting. Enumerate every endpoint and method, because APIs commonly expose verbs the documentation omits. Then test authentication and authorisation on every endpoint independently, since the common failure is a well-protected front door and unprotected individual routes. Then the data: mass assignment, excessive data exposure where the API returns more than the client displays, and injection through parameters. Rate limiting and resource consumption matter too, because APIs are often deployed without either.

What they are assessing

Whether you treat APIs as a distinct surface rather than a web app without a browser.

What is SSRF, and why is it dangerous in cloud environments?midtechnicalCommonly asked

Server-Side Request Forgery is a vulnerability where an application can be tricked into making HTTP or other network requests to destinations the attacker chooses, using the *server’s* network position and privileges rather than the attacker’s. It is dangerous generally because the server can usually reach internal systems an external attacker cannot — internal admin panels, databases, other services behind the firewall — so SSRF turns the vulnerable server into a proxy into the internal network. It became genuinely critical with the move to cloud because of one specific detail: cloud instances expose an instance metadata service on a link-local address (the well-known 169.254.169.254), and on an unprotected configuration that endpoint will hand back temporary credentials for the IAM role attached to the instance. So an SSRF that can reach the metadata endpoint escalates in a single step from “I can make the server fetch a URL” to “I now hold the cloud credentials of that server’s role,” which frequently means broad access to the account. That one-step jump from a request-forgery bug to full credential compromise is why SSRF has driven several very large cloud breaches, the Capital One breach being the canonical example. The senior-level detail worth adding: the defence is IMDSv2 (which requires a session token and blocks the naive SSRF path), so a strong answer pairs the attack with knowing that the mitigation is enforcing IMDSv2 and locking down egress, not just ‘validating URLs’.

What they are assessing

Whether you understand why one vulnerability class became disproportionately severe in cloud.

How would you test for SQL injection safely on a production system?midtechnicalCommonly asked

Confirm without extracting — that discipline is the whole answer. Establish that your input actually reaches the query by observing behavioural differences rather than pulling data: a boolean condition that reliably changes the response (true vs false returning different pages), or a time-based test where you inject a controllable delay and watch the response time move with it. To prove real impact, retrieve something harmless and non-personal — a database version banner, the current database user, a row count — which is conclusive evidence that injection exists and is exploitable, without ever touching customer records. What you emphatically do *not* do on a production system is dump tables to ‘demonstrate severity,’ because extracting real personal data is you *causing* the exact breach you were hired to find and prevent — it can be a notifiable data breach in its own right, and ‘but I was testing’ is not a defence that makes the client’s regulatory obligation disappear. If proof of actual data access is genuinely required for the report, you agree it in writing in the rules of engagement first and use seeded, non-real test records planted for the purpose. The maturity being tested here is that a good tester proves the vulnerability and its impact while causing zero harm — demonstrating you *could* read the customer table is the finding; actually reading it is an incident.

What they are assessing

Professional restraint, which is marked as heavily as technical skill in UK consultancies.

What is out-of-band exploitation, and when do you need it?seniortechnical

Out-of-band exploitation means triggering the target to make a network connection *out* to infrastructure you control, so you can confirm a vulnerability whose effect is otherwise completely invisible in the application’s own response. You need it whenever exploitation is ‘blind’: SQL injection where no query output is ever reflected back to you, SSRF where you can make the server request a URL but never see the result, or deserialisation where successful code execution changes nothing visible on screen. In all these cases the application gives you no in-band signal that the attack worked. The out-of-band technique solves this: you make the target resolve a unique DNS subdomain or send an HTTP request to a listener you control (Burp Collaborator is the standard tool), and the moment that callback lands on your infrastructure, you have proof of execution that nothing in the application response could have given you. Two senior points. First, DNS-based callbacks often succeed where HTTP ones are blocked, because egress firewalls that stop outbound web traffic frequently still allow DNS resolution — so DNS is the more reliable channel in locked-down environments. Second, and it’s a scope-discipline point testers forget: the callback infrastructure is *yours* and sits outside the client’s estate, so its use has to be agreed in the rules of engagement rather than assumed to be fine — exfiltrating even a DNS lookup to third-party infrastructure without authorisation is the kind of thing that turns a clean engagement into an awkward conversation.

What they are assessing

Technique depth plus the scope awareness that should accompany it.

How would you test file upload functionality?midtechnical

Establish what the application accepts and, crucially, *how* it validates — because the whole attack surface is in the gap between what the app thinks it's accepting and what it actually accepts. Try extensions that execute on the server platform (.php, .aspx, .jsp depending on the stack), content types that contradict the extension (a PHP payload with an image/jpeg Content-Type), double extensions (shell.php.jpg, which some misconfigured servers execute), and null-byte or case tricks against naive filters. Check whether validation is client-side only — if the only check is JavaScript in the browser, it's trivially bypassed by intercepting and modifying the request in a proxy, so always test the server's behaviour directly rather than trusting the front end. Then find where the file lands, because that determines severity: if an uploaded file is stored inside the web root and is directly reachable by URL, an uploaded executable becomes remote code execution — the highest-impact outcome and the one to prove (safely, with a benign proof-of-execution rather than a real web shell). Also test the non-execution paths that people forget: path traversal in the filename (../../ to write outside the intended directory), archive-extraction issues (a zip that unpacks over existing files, 'zip slip'), vulnerabilities in image-parsing libraries triggered by a malformed image, and files large enough to cause denial of service. A strong answer frames it as a chain: what's accepted, how it's validated, where it's stored, and whether it's reachable — because a file upload is only critical when all four line up.

What they are assessing

Coverage of a high-severity feature that juniors often test superficially.

What is insecure deserialisation, and how would you approach testing for it?seniortechnical

Insecure deserialisation is where an application takes data it received from outside and reconstructs (‘deserialises’) it back into in-memory objects, and an attacker who controls that incoming data can influence what gets constructed — or what code runs *during* the reconstruction. It matters more than most vulnerability classes because successful exploitation is typically full remote code execution rather than something milder like information disclosure, so it tends to be a critical-severity finding when it lands. Approaching testing for it: start by identifying serialised data in transit, which means recognising the format signatures — Java serialised objects have a distinctive header, PHP and .NET and Python pickle each have recognisable patterns, and they turn up in cookies, hidden form fields, view-state parameters, and custom headers. Once you’ve spotted serialised data, determine the framework and, crucially, whether known ‘gadget chains’ exist for the libraries the application uses — tools like ysoserial generate these ready-made payloads for common Java libraries, which is often how exploitation actually happens in practice. The honest, senior caveat is that insecure deserialisation is one of the *harder* classes to find blind, precisely because the effect is often invisible in the response (which is exactly when you reach for the out-of-band callback technique to confirm execution). That difficulty is why it’s frequently caught through source-code review or by recognising the format signature in traffic, rather than by black-box fuzzing — a point worth making because it shows you understand how it’s found in reality, not just what it is.

What they are assessing

Whether you can discuss a harder vulnerability class credibly.

Walk me through how you would test an internal network.midtechnicalCommonly asked

Start passive: listen before you touch anything. Broadcast and multicast traffic on a corporate network volunteers a great deal — hostnames, name resolution requests, sometimes credentials. Then discovery: what is alive, what services are exposed, what the domain structure looks like. Then the classic quick wins — default credentials, unauthenticated services, unpatched systems, SMB signing and relay opportunities. Then Active Directory specifically, because in a Windows estate that is where the path to domain compromise lives. Throughout, prioritise by what leads to privilege rather than by what is simply vulnerable.

What they are assessing

Structured internal methodology rather than a tool sequence.

What are the common paths to domain admin?midtechnicalCommonly asked

Most real compromises reach domain admin through a small number of recurring, unglamorous routes rather than exotic zero-days — and naming them shows you understand how Active Directory actually falls. The common paths: credential harvesting from a compromised host followed by reuse, especially where the same local administrator password is set identically across the estate, so one cracked machine unlocks hundreds (the reason Microsoft’s LAPS, which randomises local admin passwords, exists). Kerberoasting a service account that has a weak, crackable password *and* privileged group membership — the combination is what makes it lethal. Abusing excessive Active Directory permissions: rights delegated to a group years ago, never reviewed, that turn out to chain into domain admin — this is exactly what tools like BloodHound surface by mapping the attack paths graphically, and it’s frequently the fastest route in a mature domain. NTLM relay attacks where SMB signing is not enforced, letting you relay an authentication to a system that will accept it. And unconstrained delegation, where a compromised server configured for it will capture the Kerberos tickets of anyone who connects to it, including, if you can coerce them, a domain admin. The senior framing worth adding: notice that almost none of these are software vulnerabilities to be patched — they are *misconfigurations and hygiene failures*, which is why domain-admin compromise is usually a configuration-review and tiering problem rather than a patching one, and why the remediation is things like LAPS, tiered admin, and signing enforcement rather than a Windows update.

What they are assessing

Whether you know the realistic attack paths rather than a list of techniques.

What is Kerberoasting, and how would you exploit it?midtechnicalCommonly asked

Kerberoasting exploits a design feature of Kerberos: any authenticated domain user can request a service ticket (a TGS) for any account that has a Service Principal Name registered, and part of that ticket is encrypted with the target service account’s password hash. So the attack is: request the ticket as a normal low-privileged user, take it away, and crack it entirely *offline* — no further interaction with the domain, no failed logons, nothing noisy. That offline aspect is what makes it so attractive to attackers and so important for defenders to understand: the only on-network footprint is a single, legitimate-looking service-ticket request that blends into normal traffic, and everything after that happens on the attacker’s own hardware where no detection can see it. The value of the attack depends *entirely* on the target account, which is the key point for both exploitation and reporting: a service account with a strong, long, random password is effectively immune because it won’t crack in reasonable time, whereas a service account with a twelve-year-old human-chosen password *and* Domain Admin membership is a straight, quiet path to full compromise. When you report it, that context is the finding — name the specific vulnerable accounts, their privilege levels, and their password ages, because ‘Kerberoasting is possible’ as a generic statement is far less actionable than ‘this Domain Admin service account has a 2013-vintage password crackable in hours.’ The remediation to recommend: long random passwords or group Managed Service Accounts for service accounts, and stripping unnecessary privilege from them.

What they are assessing

Practical exploitation understanding plus reporting that drives remediation.

What is Pass-the-Hash, and what does it tell you about a network?midtechnicalCommonly asked

Authenticating with a stolen NTLM hash rather than a plaintext password, exploiting the fact that for NTLM authentication the hash *itself* is sufficient proof of identity — you never need to crack it back to the password, you just present it. Mechanically, an attacker who dumps hashes from a compromised machine's memory (with a tool like Mimikatz) can then authenticate to other systems as that user by passing the hash directly. But the mechanics are the less interesting half; what its success *tells you about a network* is what belongs in the report and what a senior tester leads with. A successful Pass-the-Hash means privileged credentials are being exposed on machines that should not hold them, and that administrative tiering is either absent or not enforced. The finding is rarely 'hashes can be passed' — that is simply how NTLM works and cannot be 'patched.' The real, actionable finding is the *why*: a domain administrator logged into an ordinary workstation, leaving their credential material in that workstation's memory for anyone who later compromises it, and the estate has no tiering model preventing high-privilege accounts from touching low-trust machines. That reframing matters because it points at the actual fix — Microsoft's tiered administration model, so domain admins only ever log into domain controllers and secured admin workstations; Credential Guard to protect hashes in memory; and stripping unnecessary privilege — rather than the futile goal of 'stopping Pass-the-Hash,' which is like trying to stop the letter 'e.' The technique is a symptom; the credential-hygiene and tiering failure is the disease, and the disease is what you report.

What they are assessing

Whether you can turn a technique into an architectural finding, which is what senior testers do.

How would you approach a wireless assessment?seniortechnical

Establish what is deployed and what should be: enterprise authentication, pre-shared key networks, guest networks, and anything unauthorised. For pre-shared key networks, capture a handshake and test the passphrase offline. For enterprise networks, the interesting weakness is usually certificate validation — if clients do not properly validate the authentication server, an evil twin can harvest credentials. Check guest network isolation, since guest networks frequently reach internal resources they should not. And look for rogue access points, including well-meaning ones staff have plugged in themselves.

What they are assessing

Whether you can cover a specialism many testers only know superficially.

How would you test a mobile application?midtechnical

Three surfaces, and juniors usually test only one. The client itself: how it stores data locally, whether anything sensitive lands in plaintext, and how it handles secrets embedded in the package. The transport: whether certificate pinning is implemented and whether it can be bypassed, and whether the app falls back to insecure connections. And the backend API, which is usually where the real findings are, because mobile developers sometimes assume the app is the only client. Testing typically requires a rooted or jailbroken device, or an emulator, plus an interception proxy configured to handle pinning.

What they are assessing

Whether you understand mobile as three surfaces rather than one.

How would you plan a phishing engagement?midtechnical

Objectives first: are you measuring click rates, harvesting credentials, or achieving access as part of a wider engagement? Each implies a different design. Then authorisation in writing, including who is in scope and explicitly who is not — targeting individuals can have employment and wellbeing consequences, which is why scope matters more here than almost anywhere. Then infrastructure: domains, mail configuration, and hosting that will actually deliver. Then the pretext, which should be plausible for that organisation rather than generically clever. And agree in advance how findings will be reported, because naming individuals is usually wrong.

What they are assessing

Ethical and operational planning, not just pretext writing.

How would you test physical security?seniortechnical

With unusually careful authorisation, because this is the engagement type most likely to end with a tester talking to the police. You need a written authorisation letter carried on the person, a named contact reachable at any hour who can confirm the engagement, and clarity on which buildings, which hours, and which techniques are permitted. Then the assessment itself: perimeter, access control, tailgating, reception process, whether challenge culture exists, and what is reachable once inside — unlocked workstations, network ports, documents, server room access. The professional standard is that you stop and identify yourself when challenged.

What they are assessing

Whether you understand the legal exposure and control requirements of physical work.

How do you rate the severity of a finding?midtechnicalCommonly asked

Most reports use CVSS as a common baseline, and you should — it gives comparability across findings and clients, and many clients contractually expect it. But the honest caveat, and the thing that separates a mature tester from a junior one in an interview, is that CVSS base scores *deliberately* exclude environmental context. The base score rates the intrinsic technical severity of a flaw in the abstract, which means the same vulnerability scores *identically* whether it’s on an isolated, non-sensitive development box or on the internet-facing system that processes card payments — and that’s obviously wrong for the client’s actual prioritisation, because those two situations demand completely different urgency. So the mature approach is to use CVSS as the technical baseline and then adjust for the things it omits: real-world exposure (internet-facing vs internal vs air-gapped), the sensitivity of the data at risk, and the concrete business impact if it’s exploited — and, importantly, to *state the adjustment reasoning* so the client can see why you rated it as you did rather than just trusting a number. CVSS has a Temporal and Environmental metric set intended for exactly this, but the principle matters more than the mechanism. The junior tell the interviewer is listening for is mechanical CVSS with no contextual adjustment — a report that dutifully scores everything and ranks a high-CVSS bug on a dev server above a medium-CVSS bug on the payments platform, because it has confused technical severity with business risk.

What they are assessing

Whether you can use CVSS while understanding its limits.

How would you handle a client disputing a finding?midcompetencyCommonly asked

Separate the observation from the rating, because usually the disagreement is about severity rather than existence, and that is a far more productive conversation. Go back to the evidence: show the request and response, or reproduce it. If the client has context you lacked — a compensating control you could not see, a system due for decommission — then adjust the rating and say so, because being correctable builds more trust than being immovable. If you still disagree, record both positions in the report. What you never do is inflate or deflate a severity to manage a relationship, because the moment ratings become negotiable, every future finding is discounted.

What they are assessing

Professional handling under commercial pressure, and ratings integrity.

How would you write a finding so it actually gets fixed?midtechnicalCommonly asked

Write for the person who has to do the work — the developer, not the auditor. A finding that gets fixed has a specific shape. State what the issue is in one clear sentence up front. Then give reproduction steps precise enough that the developer can see it happen themselves, because a bug they can reproduce is a bug they believe, and one they can’t is one they’ll dispute. Then the evidence. Then remediation that is *specific to their stack*, not generic advice — and this is where most findings fail. ‘Sanitise user input’ is not remediation; it’s a slogan that tells the developer nothing about what to actually change. ‘Use parameterised queries in this data-access layer — here’s the exact call, rewritten’ is remediation, because it can be implemented without further research. Explain the impact in terms that justify the effort, because developers are prioritising your finding against feature work with real deadlines, and ‘this is critical’ means nothing next to ‘this lets an unauthenticated user read every customer’s records.’ And avoid condescension entirely — a finding that reads as a rebuke to the person who wrote the code gets argued with, escalated, and slow-walked, whereas one that reads as a colleague helpfully pointing at a problem gets fixed. The whole craft is remembering that a finding’s job is not to demonstrate that you found something clever; it’s to cause a change in someone else’s codebase, and everything about how you write it should serve that.

What they are assessing

Whether you write to change outcomes rather than to demonstrate skill.

What would you include in an executive summary?midtechnicalCommonly asked

What the test covered, what the overall security posture looks like, what the most significant risks are in business terms, and what should be done first. No payload syntax, no CVE numbers, no tool names. The audience is deciding whether to fund remediation and how urgently, so the summary should answer "how exposed are we and what does it cost to fix" rather than describing what you did. One page is usually right. The commonest failure is an executive summary that is simply a shorter technical section, which leaves the decision-maker with information they cannot act on.

What they are assessing

Whether you can write for a non-technical decision-maker.

You find nothing significant on an engagement. What do you report?midscenarioCommonly asked

Report exactly that — honestly, and with the coverage stated so the client can correctly interpret what a clean result means. A clean report is a completely legitimate outcome and a client who has a well-secured system is *entitled* to be told so; pretending otherwise is where testers lose their integrity. But ‘nothing significant’ is only meaningful alongside what was and wasn’t tested: the scope covered, the constraints that applied (time-boxing, an environment that kept falling over, credentials that arrived late), and an explicit statement of what the absence of findings does and does not prove — ‘no critical vulnerabilities were identified in the tested scope within the engagement window’ is honest; ‘the system is secure’ is not. What you must *never* do is inflate trivial observations into findings to justify the fee — dressing up ‘the server returns a version header’ as a vulnerability to make the report look busy. It’s a slow-acting poison: it destroys your credibility with the technical people who read it, and it trains the client to discount *all* your severities, so that when you later report something genuinely critical they’ve already learned to shrug. Two things worth doing instead: note the *good* practice you observed, which is genuinely useful information a client almost never gets and which builds trust, and be honest in the debrief that a clean result within a limited scope is not the same as a guarantee. A confident, well-scoped clean report is a sign of a mature tester; a padded one is a sign of an insecure one.

What they are assessing

Integrity under commercial pressure, which UK consultancies screen for deliberately.

You are running behind and will not finish the agreed scope. What do you do?midscenario

Communicate early — the moment you realise you won’t finish, not at the end when it’s a fait accompli. The instinct under time pressure is to quietly cut corners, skim the remaining scope, and present it as complete; that instinct is the actual failure here, because a test presented as full coverage when it wasn’t gives the client false assurance about systems that were never really looked at, and that false assurance is worse than no test at all — it stops them worrying about something they should. So the honest move is: as soon as it’s clear the timeline is slipping, tell the engagement lead and the client, and re-prioritise together. Focus the remaining time on the highest-risk parts of the scope — the internet-facing, the sensitive, the business-critical — and consciously defer the lower-risk remainder, so that what you *do* test is the right stuff. Then, in the report, state plainly what was and wasn’t covered, so the client knows precisely where they have assurance and where they don’t. Depending on the contract, the deferred scope might become a follow-up engagement or an extension. The judgement being tested is whether you’ll protect the client’s accurate understanding of their own risk over your own discomfort at admitting you ran out of time — a tester who says ‘I covered the critical systems fully and here’s exactly what I didn’t reach’ is far more valuable than one who claims full coverage they didn’t deliver.

What they are assessing

Commercial honesty and communication discipline.

You accidentally cause an outage during testing. What do you do?midscenarioCommonly asked

Stop, and tell them immediately — that sequence, in that order. The instant an outage happens, halt the activity that caused it so you don’t make it worse, then contact your client point of contact straight away; do not keep testing and do not wait to see if it recovers on its own while saying nothing. Outages happen in testing even to careful testers — fragile legacy systems fall over, a scan overwhelms something under-resourced — so the outage itself is not automatically negligence, but *how you handle it* is what defines you. Be completely transparent about what you were doing when it happened, because that information is exactly what the client needs to recover quickly, and a tester who obscures their actions to avoid blame directly extends the outage they caused. Help with recovery to the extent you usefully can, and afterwards adjust the approach — more conservative settings, testing in a window agreed with them, avoiding the fragile component or handling it more gently. Two points that show maturity: this is precisely why rules of engagement should establish emergency contacts and stop conditions *before* testing starts, so there’s a plan for this moment rather than improvisation; and honesty here protects the relationship far more than it damages it — clients forgive an accident handled with transparency, but they do not forgive discovering that the tester knew and hid it. The failure mode being tested is the tester who, out of embarrassment, goes quiet and lets the client burn time diagnosing a problem the tester could have explained in thirty seconds.

What they are assessing

Whether you own mistakes immediately, which is heavily weighted in consultancy hiring.

What is your approach to using automated tools?entrytechnicalCommonly asked

As accelerators for coverage, not substitutes for testing. Scanners find the known and the mechanical efficiently, which frees human time for the things they cannot do — access control, business logic, chained issues. The disciplines that matter: understand what a tool actually does before pointing it at a client system, verify every finding before it reaches a report, and never let tool output become the report. A tester whose deliverable is exported scanner results has sold a scan at penetration test prices, which is the single most common complaint clients have about the industry.

What they are assessing

Balanced relationship with tooling, neither dismissive nor dependent.

What tools would you use for a web application test, and why?entrytechnicalCommonly asked

An intercepting proxy throughout — Burp Suite is the industry standard and most of the work happens there, with ZAP as the open-source alternative. Beyond that, tooling is chosen per task rather than run as a suite: content discovery for unlinked endpoints, specific tools for particular classes where they genuinely add speed, and manual testing for anything requiring reasoning about what the application does. The answer that scores badly is a long list of tool names with no account of what each is for, because it suggests the candidate runs tools rather than tests applications.

What they are assessing

Practical tooling knowledge with reasoning attached.

How would you approach source code review?seniortechnical

Start where the risk concentrates rather than reading linearly. Find the entry points — anywhere untrusted input enters — then trace it forward to the dangerous sinks: query construction, command execution, deserialisation, file operations, rendering. Then check authentication and authorisation logic specifically, because that is where the highest-impact flaws hide and where automated tooling is weakest. Automated scanning helps with breadth and produces substantial false positives, so treat it as a lead generator. The advantage over black-box testing is that you can see the logic rather than infer it, which finds classes of flaw that testing from outside would never surface.

What they are assessing

Whether you can review code methodically rather than reading hopefully.

How do you keep your technical skills current?entrycompetencyCommonly asked

Deliberately, and in a way you can evidence. The routes that work: a home lab where you can break things safely, hands-on platforms and capture-the-flag exercises, following a small number of researchers whose work you actually read, and reproducing published techniques rather than only reading about them. The answer that fails is a list of subscriptions with nothing behind it, because the follow-up is always "what have you looked at recently?" and a memorised list collapses immediately. One technique you have actually reproduced beats fifty you have read summaries of.

What they are assessing

Genuine currency, tested by the inevitable follow-up.

What would you do if you found a vulnerability outside the agreed scope?entryscenarioCommonly asked

Stop, and report it — do not touch it. Finding a vulnerability outside the agreed scope is a real test of professional discipline, because the technical temptation is obvious: it’s *right there*, and exploring it feels like doing a thorough job. But acting on it is exactly the wrong move, and understanding why is the point. Testing outside the authorised scope isn’t just a contractual breach — depending on what the out-of-scope system is, it can cross into unauthorised access, and the written authorisation that makes your testing legal (rather than a Computer Misuse Act offence) only covers what was actually agreed. So the correct action is: note what you found and how you found it, do not probe or exploit it further, and report it to the client through the proper channel so *they* can decide whether to bring it into scope, commission a separate piece of work, or handle it themselves. Often they’ll be glad you spotted it and will authorise a look; sometimes the out-of-scope system belongs to a third party (a shared host, a supplier’s infrastructure) and testing it would have created a genuine legal problem for everyone. The maturity being assessed is that a professional tester understands scope as a legal and ethical boundary, not a bureaucratic inconvenience to route around — the discipline to find something interesting, leave it alone, and hand the decision to the person entitled to make it is precisely what makes a client able to trust you with access in the first place.

What they are assessing

Scope discipline, which UK consultancies treat as close to disqualifying if absent.

What is in scope and out of scope in an engagement, and why does it matter?entrytechnicalCommonly asked

Scope defines what you are authorised to test — systems, addresses, applications, techniques, timing. Rules of engagement wrap it with the operational detail: escalation contacts, what to do on a critical finding or suspected existing compromise, how discovered data is handled, and limits on disruptive techniques. It matters because authorisation is the only difference between penetration testing and a criminal offence. Out of scope is not a suggestion: it may be a third party’s system, shared infrastructure, or something in a legal state you do not know about.

What they are assessing

Whether scope is understood as the ethical foundation rather than paperwork.

What is responsible disclosure, and how does it work?midtechnical

Reporting a vulnerability privately to the party who can fix it, allowing a reasonable period for remediation before any public disclosure. In practice: report through their published security contact, provide enough detail to reproduce, agree a timeline, and coordinate any publication. Where no contact exists, national CERTs can act as intermediaries. The tension is real — vendors sometimes stall indefinitely, and disclosure is the only remaining pressure — but the professional position is to exhaust coordination first. Note also that finding the vulnerability may itself require testing you were not authorised to perform, which is a separate legal question.

What they are assessing

Ethical framework and awareness of the legal complication in the UK.

What is a bug bounty programme, and how does it differ from a penetration test?midtechnical

A bug bounty invites a large pool of researchers to find vulnerabilities continuously, paying per valid finding. A penetration test buys defined coverage from named testers within a fixed period. They solve different problems: bounties give breadth, continuous attention and only pay for results, but coverage is unpredictable and researchers gravitate toward the easily monetised. A test gives assured coverage, a report you can hand an auditor, and testing of things nobody would hunt voluntarily. Mature organisations run both, and treating a bounty as a substitute for a test usually reflects a budget decision rather than a security one.

What they are assessing

Commercial understanding of the offensive security market.

What is the difference between a vulnerability scan and a penetration test?entrytechnicalCommonly asked

A vulnerability scan is automated and broad: a tool checks systems against a database of known issues — missing patches, default credentials, known-vulnerable software versions — and produces a list, fast and cheaply, across a large estate. A penetration test is a human-led assessment that goes further in every dimension that matters: a tester validates whether findings are *actually* exploitable (filtering out the false positives scanners are notorious for), chains multiple lower-severity issues into a real attack path the way an actual adversary would, tests business logic that no scanner understands, and interprets impact in the context of the specific organisation. The distinction people miss, and the one worth making, is that they answer different questions: a scan answers ‘which known issues are present on these systems?’ while a pentest answers ‘what could an attacker actually achieve here?’ — and those are not the same question, because a scanner reports a hundred findings without knowing which one is the real way in, and a tester’s value is precisely in knowing. Both have their place: scanning is how you maintain continuous baseline hygiene between tests, cheaply and often; penetration testing is the periodic, deeper, human assessment that finds what scanners structurally cannot. A strong answer resists the common trap of implying a scan is a cheap pentest — they’re complementary tools for different jobs, and a mature security programme uses both rather than mistaking one for the other.

What they are assessing

Whether you respect both activities rather than sneering at automation.

How would you explain a critical finding to a non-technical client on a call?midcompetencyCommonly asked

Lead with the impact in their terms, not the mechanism — the same translation skill that matters at board level, applied at the pointed moment of delivering bad news. A non-technical stakeholder does not need to understand SQL injection, deserialisation, or the exploit chain; they need to understand three things clearly: what could happen, how bad it would be for *them* specifically, and what you’re recommending they do about it. So you translate: not ‘there’s an unauthenticated SQL injection in the customer portal,’ but ‘an attacker on the internet, with no account and no password, could read your entire customer database — every record — and we should treat fixing it as urgent.’ Anchor the severity in a consequence they already care about (customer data exposure, a reportable breach, service down, regulatory exposure) rather than a CVSS number that means nothing to them. Give them a clear recommended action and a sense of urgency proportionate to the real risk, so they can make a decision rather than just absorb alarming information. And read the room: the goal is informed action, not frightening them into paralysis or, worse, into shooting the messenger — a stakeholder who feels blamed or bewildered disengages. The skill being tested is whether you can make a genuinely critical technical finding *land* with someone who can’t evaluate it technically, so that it drives the fix rather than a panic or a shrug — which is often the difference between a finding that gets remediated and one that gets filed.

What they are assessing

Client-facing communication under pressure, which consultancies weight heavily.

How would you prepare for a technical interview involving a practical assessment?entrycompetencyCommonly asked

Expect to be watched thinking rather than tested on recall. Practise narrating your reasoning aloud, because assessors mark method and most candidates go silent when concentrating. Rehearse the fundamentals until they are automatic — enumeration, common web classes, basic privilege escalation — so cognitive effort goes to the unusual part. Know your tooling well enough not to fumble syntax under observation. And prepare for the moment you get stuck, because it will happen: say what you have ruled out and what you would try next, which scores far better than silence or guessing.

What they are assessing

Self-awareness about how practical assessments are actually marked.

What is privilege escalation, and how do you approach it on Linux?midtechnicalCommonly asked

Moving from the access you have to more than you should — from a normal user to root. On Linux the routine checks form a recognisable methodology, and naming it shows you’ve actually done this rather than read about it. Kernel version against known local-privilege-escalation exploits (useful but riskier, since kernel exploits can crash the box). SUID and SGID binaries — files that run with the owner’s privileges — especially unusual ones or known-abusable binaries (GTFOBins is the reference for which standard binaries can be turned into privilege escalation). Sudo rights: what the current user can run with sudo, since a single mis-granted sudo entry on an editor, an interpreter, or a binary that can spawn a shell is often the whole path. Writable files that root uses — cron jobs owned by root but editable by you, writable scripts in root’s PATH, world-writable service files. Misconfigured file permissions on sensitive things like /etc/passwd or private keys. Credentials left in config files, history files, and environment variables. Capabilities set on binaries, a subtler modern equivalent of SUID. The senior framing worth adding: the great majority of real Linux privilege escalation is *misconfiguration* — a sudo rule, a writable cron, a leaked credential — not kernel exploitation, which is why tools like LinPEAS automate the enumeration of exactly these checks, and why the finding you report is usually ‘this specific misconfiguration grants root’ rather than ‘the kernel is old.’ Report the concrete path you used, not a generic ‘privilege escalation possible.’

What they are assessing

Practical escalation knowledge with operational caution attached.

What is privilege escalation, and how do you approach it on Windows?midtechnicalCommonly asked

Moving from the access you have to more than you should — the routine paths on Windows differ from Linux and reflect the different OS design. Service misconfigurations are the classic starting point: unquoted service paths (where a space in an unquoted path lets you plant a binary Windows will execute), weak service permissions (a service you can reconfigure to run your binary), and writable service executables. Registry autoruns and scheduled tasks pointing at files or locations you can write to. ‘AlwaysInstallElevated,’ a policy misconfiguration that lets any user install an MSI as SYSTEM. Stored credentials — in the Credential Manager, in unattended-install files, in the registry, in scripts. Token impersonation and the family of ‘Potato’ techniques that abuse Windows privileges (like SeImpersonate, often held by service accounts) to escalate to SYSTEM, which is why compromising a web or database service account is frequently a stepping stone rather than an endpoint. DLL hijacking where a privileged process loads a DLL from a location you control. And missing patches for known local exploits. As on Linux, the honest senior framing is that most real Windows escalation is misconfiguration and privilege abuse rather than kernel exploitation, which is why enumeration tools like WinPEAS and PowerUp exist to surface exactly these paths automatically — and why the remediation you recommend is usually service-permission hygiene, credential cleanup, and disabling AlwaysInstallElevated rather than ‘patch the kernel.’ Report the specific path you took to SYSTEM, with the misconfiguration that enabled it, because that’s what gets fixed.

What they are assessing

Windows-specific depth, which matters most in UK enterprise estates.

Cloud Security

44
What is the shared responsibility model, and where do organisations get it wrong?entrytechnicalCommonly asked

The division of security duties between the cloud provider and the customer: the provider secures the underlying infrastructure, and the customer secures what they put on it. The line moves by service type — for infrastructure-as-a-service the customer owns almost everything above the hypervisor, while for software-as-a-service they own little beyond their data and access. Where organisations get it wrong is assuming the provider covers more than it does, particularly configuration and identity: the provider will keep the platform running securely and will happily let you expose your own storage to the internet. Most cloud breaches live squarely on the customer’s side of the line.

What they are assessing

Whether you understand the model as a moving line rather than a fixed split, and where customer failures cluster.

Why is identity the new perimeter in cloud?entrytechnicalCommonly asked

Because in cloud there is no network edge to defend in the traditional sense — resources are reached over the internet, authenticated by identity, and the thing that decides whether a request succeeds is almost always a credential and a permission rather than a network position. An attacker with a valid access key is inside, wherever they physically are. This is why the highest-impact cloud attacks are identity attacks: stolen keys, over-permissioned roles, compromised service principals. The defensive consequence is that identity controls — strong authentication, least privilege, conditional access, credential hygiene — do the work that firewalls used to.

What they are assessing

Whether you grasp the fundamental shift from network-centric to identity-centric security.

What is the first thing you would look at in an unfamiliar cloud environment?midtechnicalCommonly asked

Identity and exposure, because those are where the fastest-moving, highest-consequence risk lives — and starting there rather than with an inventory is itself the thing that signals seniority. On identity: who has privileged access (and is it far more people than anyone realises), whether there are unused or over-permissioned roles accumulated over time, whether multi-factor is genuinely enforced rather than merely available, and whether long-lived access keys are sitting around that could be leaked or stolen. On exposure: what is reachable from the internet that should not be — publicly readable storage, management ports open to the world, databases with public endpoints, admin interfaces without IP restriction. A cloud posture tool (CSPM) gives that picture in minutes if one is deployed, and its absence is itself a finding. The reason to lead with identity and exposure rather than a methodical asset inventory is that you’re thinking like the attacker: they don’t enumerate everything you own, they look for the one public bucket and the one over-privileged role that gets them in, so the first question isn’t ‘what is here?’ but ‘what would an attacker reach first, and what could they do once they did?’ The senior instinct being tested is prioritising the blast-radius questions — who’s over-privileged, what’s exposed — over the comforting but slow completeness of a full inventory.

What they are assessing

Whether you prioritise by attacker-reachable risk rather than trying to boil the ocean.

How does networking differ in the cloud from on-premises?midtechnical

The primitives look familiar but the model is inverted. On-premises you defend a perimeter and largely trust the inside; in cloud you build software-defined networks where segmentation is cheap, everything is potentially reachable unless you say otherwise, and the meaningful boundary is often identity rather than subnet. Security groups and network ACLs replace physical firewalls, are defined as code, and can be changed in seconds — for better and worse. The biggest conceptual shift is that network controls are necessary but no longer sufficient, because a correctly firewalled resource is still fully exposed if its access keys leak.

What they are assessing

Whether you understand cloud networking as a different model rather than the same one virtualised.

What is the difference between a security group and a network ACL?midtechnical

Both filter traffic, but they operate differently. A security group is attached to a resource and is stateful — if you allow inbound traffic, the response is automatically allowed out, and it only has allow rules. A network ACL operates at the subnet level, is stateless — you must permit both directions explicitly — and supports both allow and deny rules, processed in order. In practice security groups do most of the work because statefulness makes them simpler and less error-prone, while network ACLs provide a coarser subnet-level backstop, often used for broad denies. Confusing the two, particularly forgetting NACL statelessness, causes hard-to-debug connectivity failures.

What they are assessing

Practical networking precision that distinguishes hands-on candidates.

What is the difference between a managed identity, a service principal and a service account?midtechnicalCommonly asked

In Azure terms: a service principal is an identity for an application or service to authenticate and be granted access; a managed identity is a special kind of service principal whose credentials are handled entirely by the platform, so there is no secret for you to store, rotate or leak. That last point is the whole value — managed identities remove the stored credential, which is one of the most common cloud attack routes. A service account is the more generic, often on-premises, term for a non-human account. The interview-worthy point is that managed identities should be the default wherever the platform supports them, precisely because they eliminate the secret.

What they are assessing

Whether you understand why managed identities are preferred, not just what they are.

Why are long-lived access keys dangerous, and what would you use instead?midtechnicalCommonly asked

Because a long-lived key is a static secret that keeps working until someone notices it has leaked, and secrets leak constantly — into code repositories, logs, configuration files, laptops. A key with no expiry that grants standing access is exactly what an attacker wants, and the dwell time between leak and discovery is often long. The alternative is to eliminate the standing secret: managed identities or IAM roles for service-to-service access, short-lived tokens obtained on demand, and federated credentials for CI/CD so pipelines authenticate without a stored key at all. Where a static key is genuinely unavoidable, it goes in a secrets manager with rotation.

What they are assessing

Whether you reach for credential elimination rather than just credential rotation.

Explain least privilege in a cloud context, and why it is so hard.midtechnicalCommonly asked

Granting each identity only the permissions it actually needs, and no more. It is the right principle and genuinely hard in cloud for structural reasons: permission systems are enormous and granular, so working out the minimal set for a workload is real effort; over-provisioning is the path of least resistance because broad access makes things work immediately; and tightening later risks breaking a production dependency nobody fully understands. The result is that most cloud identities accumulate far more permission than they use. The practical approach is to start restrictive and widen on evidence, and to use the tooling that reports actual-versus-granted permissions.

What they are assessing

Whether you understand why least privilege fails in practice, not just what it means.

What is privilege escalation in a cloud IAM context?seniortechnical

Using permissions you have to obtain permissions you should not, without exploiting any software vulnerability — the escalation is in the permission model itself. Classic routes: a permission to modify IAM policies lets an identity grant itself more; the ability to pass a role to a service lets a low-privileged user launch a resource that runs as a high-privileged role; permission to update a function’s code lets you run arbitrary actions under that function’s identity. These are configuration flaws rather than exploits, which is what makes them easy to miss and why access-path analysis matters more than vulnerability scanning in cloud.

What they are assessing

Whether you understand that cloud escalation is usually a permissions problem, not an exploit.

How would you handle a developer needing production access just for today?midscenarioCommonly asked

Give it, but through a mechanism rather than a standing grant. The right answer is just-in-time elevation: the developer requests access, it is approved, granted for a bounded window, logged, and automatically revoked — so the access exists exactly as long as the need and leaves an audit trail. Defaulting to yes with a permanent grant is how estates accumulate standing privilege; defaulting to no drives people to work around security, which is worse. The mechanism resolves the tension: the developer gets what they need, and the organisation does not carry the permission afterwards. If this keeps happening, that is a signal the process needs fixing, not the person.

What they are assessing

Whether you reach for a mechanism rather than a binary yes or no.

What is CSPM, and what problem does it actually solve?midtechnicalCommonly asked

Cloud Security Posture Management: tooling that continuously checks cloud configuration against security best practice and flags misconfigurations — public storage, unencrypted databases, over-permissive rules, missing logging. The problem it solves is scale and drift. Cloud configuration changes constantly across hundreds of resources, and misconfiguration is the dominant cause of cloud breaches, so a point-in-time manual review is obsolete the moment it finishes. CSPM makes posture continuous and visible. Its limit, worth naming, is that it finds known misconfigurations against a ruleset — it does not understand your specific business context, so it needs tuning to avoid drowning teams in low-value findings.

What they are assessing

Whether you can articulate the problem and the limitation, not just expand the acronym.

What is CIEM, and how does it differ from CSPM?seniortechnical

Cloud Infrastructure Entitlement Management focuses specifically on identities and permissions: who can do what, where the excessive entitlements are, and what the escalation paths look like. CSPM checks resource configuration; CIEM checks the permission graph. The distinction matters because cloud IAM is where the highest-impact risk concentrates and where least privilege is hardest, and configuration tooling largely misses it. A good CIEM tool answers questions like which identities could reach this sensitive resource by chaining permissions, which is exactly the analysis a manual review cannot do at cloud scale. The two are complementary and increasingly sold together as CNAPP.

What they are assessing

Whether your tooling knowledge extends to the identity-specific layer where cloud risk concentrates.

What is CNAPP?midtechnical

Cloud Native Application Protection Platform: an umbrella category bundling the previously separate cloud security tools — posture management, workload protection, entitlement management, and often infrastructure-as-code and container scanning — into one platform. The rationale is that these functions overlap and produce disconnected findings when run as separate tools, so a single platform can correlate them: an exposed workload, running vulnerable code, with an over-permissioned identity, reachable from the internet, is one prioritised finding rather than four unrelated alerts. The honest caveat is that CNAPP is a heavily marketed term, so the useful question about any specific product is which of those functions it actually does well.

What they are assessing

Current market literacy with appropriate scepticism about umbrella terms.

How would you prioritise cloud security findings when the tool reports thousands?midtechnical

By exploitability and exposure rather than by the tool’s severity label. The findings that matter are the ones that are reachable and lead somewhere: an internet-facing resource, with a real vulnerability or misconfiguration, holding or reaching sensitive data, accessible by an over-permissioned identity. That intersection is usually a short list, and it is where a real attacker would go. The volume problem is the same one SOCs face with alerts — a raw count of thousands is a triage failure waiting to happen, and the skill is converting it into a ranked queue. Tuning out the low-value rules that generate noise is part of the job, not a shortcut.

What they are assessing

Whether you can turn an unmanageable pile into a defensible queue.

What is a toxic combination in cloud risk?seniortechnical

A set of individually acceptable conditions that together create serious risk — for example, a workload that is internet-facing, running software with a known vulnerability, and attached to an identity with broad permissions. Any one of those might be tolerable; combined, they are a direct path to compromise. The concept matters because traditional scanning reports each condition separately, at modest severity, and the real risk lives in the combination. Prioritising by toxic combination is how you find the handful of genuinely dangerous exposures among thousands of individual findings, and it is the reasoning modern CNAPP tooling tries to automate.

What they are assessing

Whether you think in attack paths rather than isolated findings.

How would you secure a storage account or S3 bucket, and why does public exposure keep happening?entrytechnicalCommonly asked

Default to private, and make any public access a deliberate, blocked-by-default exception that someone has to consciously override: enforce encryption at rest and in transit, use identity-based access (IAM roles) rather than shared account keys, enable access logging so you can see who read what, and use private connectivity (VPC endpoints / private endpoints) so the data plane never touches the public internet in the first place. Most cloud providers now offer an account-level ‘block all public access’ setting that overrides individual bucket misconfigurations, and turning it on is one of the highest-value single controls available. As for why public exposure *keeps* happening despite years of headlines: the platform defaults have genuinely improved — new buckets are private by default now — so it is almost never a platform flaw. The failure is a person or a script deliberately setting public access to make something work (a developer debugging, a quick fix to unblock a demo), combined with permissions broad enough to *allow* them to do it, and no posture tooling watching to catch it afterwards. That’s the real lesson worth stating: bucket exposure is a configuration-and-governance problem, not a technology one — the fix is blocked-by-default at the account level, least-privilege so most people *can’t* make something public, and CSPM to catch the exceptions that slip through, rather than trusting everyone to configure each bucket correctly forever.

What they are assessing

Whether you can secure storage and explain the recurring human cause of exposure.

How would you manage secrets in the cloud?midtechnicalCommonly asked

The best secret is no secret at all — and leading with that principle is what distinguishes a modern answer from a dated one. Use managed identities and workload identity federation so that service-to-service authentication happens through the platform’s own identity system with no stored credential to leak, rotate, or steal; if there’s nothing to exfiltrate, an entire class of breach disappears. Where a secret genuinely must exist — a third-party API key, a legacy database password, a credential for a system outside the cloud — it goes in a dedicated secrets manager (AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, HashiCorp Vault) with access controlled by identity, automatic rotation enabled, and full audit logging of every retrieval, and it is fetched at runtime rather than ever being written into code or config. What you never do is the thing that causes most real-world credential leaks: secrets hard-coded in source, committed to a repository (where they live in the git history forever even after you delete them), baked into container images, or sitting in environment variables that get logged. Secret scanning in the CI/CD pipeline — and ideally pre-commit hooks on developer machines — catches these mistakes before they ship, which matters because the single most common way cloud credentials leak is a developer accidentally committing them to a public or later-exposed repo. The senior framing: the hierarchy is no-secret (managed identity) first, managed-secret-store second, and everything else is a finding — and a mature programme is actively driving from the second tier toward the first, eliminating stored secrets rather than just guarding them better.

What they are assessing

Whether you prioritise eliminating secrets over merely storing them well.

What is data residency, and why does it matter for UK organisations?midtechnical

Data residency is where data is physically stored and processed, which matters because law and contract frequently constrain it. For UK organisations the concerns are UK GDPR restrictions on international transfers of personal data, sector rules — parts of government and financial services expect data kept in the UK — and contractual commitments to customers. The practical implications are choosing cloud regions deliberately, understanding where a managed service actually processes and stores data including its backups and metadata, and knowing whether support access from other countries counts as a transfer. It is a governance question with real architectural consequences.

What they are assessing

UK-specific data protection awareness applied to cloud architecture.

What is sovereign cloud, and who needs it?seniortechnical

Cloud infrastructure designed to keep data and operations within a specific jurisdiction and under that jurisdiction’s control, isolated from foreign access including from the provider’s own overseas staff and from foreign legal reach. In the UK the drivers are government and highly regulated sectors handling sensitive data, and concern about foreign legislation — such as the US CLOUD Act — potentially compelling access to data held by American providers. The providers offer varying answers, from UK regions through to dedicated sovereign offerings. Whether an organisation genuinely needs it, versus UK regions with strong controls, is a risk and compliance judgement rather than a technical default.

What they are assessing

Awareness of a live UK public-sector and regulated-industry concern.

How would you approach encryption key management in the cloud?midtechnical

Decide who holds the keys, because that is the real question. Provider-managed keys are simplest and fine for most data — the provider handles rotation and storage, and you accept that they technically could access the key. Customer-managed keys, held in a cloud key management service, give you control over rotation and revocation while the provider still hosts the service. Customer-supplied or hold-your-own-key models keep the key entirely outside the provider for the most sensitive data, at real operational cost. The right choice is risk-driven: match key control to data sensitivity rather than defaulting to the most or least controlled option everywhere.

What they are assessing

Whether you can reason about key control as a spectrum matched to sensitivity.

What are the main security considerations for containers?midtechnical

Across the lifecycle. The image: build from trusted minimal bases, scan for vulnerabilities, and do not bake secrets into layers. The registry: control who can push and pull, and sign images so only trusted ones run. The runtime: run as non-root, drop unnecessary capabilities, use read-only filesystems where possible, and limit what a container can reach. And the orchestrator, usually Kubernetes, which has its own substantial attack surface. The theme is that a container is only as trustworthy as the image it came from and the privileges it runs with, and most container compromises trace back to a vulnerable image or an over-privileged runtime.

What they are assessing

Whether you cover the container lifecycle rather than a single layer.

What are the main security considerations for Kubernetes?seniortechnical

Kubernetes is powerful and insecure by default, so the considerations are broad. Control access to the API server, which is the crown jewel, with strong authentication and RBAC scoped tightly. Isolate workloads with namespaces and network policies, because flat pod networking lets a compromised pod reach everything by default. Constrain what pods can do with security contexts and admission control, preventing privileged containers and host access. Protect secrets, which Kubernetes stores base64-encoded rather than encrypted unless you configure otherwise. And secure the supply chain of images running in the cluster. The recurring theme is that the secure configuration is not the default.

What they are assessing

Whether you understand Kubernetes as insecure-by-default across several dimensions.

What are the security considerations for serverless functions?midtechnical

The attack surface shifts rather than shrinks. You no longer patch servers, which removes a class of risk, but new concerns rise in importance: each function’s permissions, because functions accumulate over-broad roles exactly like everything else in cloud IAM; the code and its dependencies, since vulnerable libraries are now your responsibility; event-source injection, because functions are triggered by data from queues, storage and APIs that may be attacker-controlled; and secrets handling, since functions often need credentials to reach other services. The common misconception is that serverless is more secure because there is no server — it is differently secure, and identity and code become the dominant concerns.

What they are assessing

Whether you understand that serverless changes the surface rather than removing it.

What security opportunities and risks come with infrastructure as code?midtechnical

The opportunity is enormous: infrastructure defined as code can be reviewed before deployment, scanned automatically, versioned so every change is auditable, and made consistent so the secure configuration is the default one that gets deployed everywhere. That turns security from a manual after-the-fact check into an automated gate. The risk is symmetrical: a mistake in a template deploys everywhere the template is used, so one insecure default becomes an estate-wide exposure, and a leaked state file or a compromised pipeline can expose or alter the whole environment. Infrastructure as code makes security either much better or much worse, and rarely leaves it unchanged.

What they are assessing

Whether you see both sides of the IaC security equation.

How would you secure a CI/CD pipeline?midtechnicalCommonly asked

Treat the pipeline as production infrastructure, because it can deploy to production and often holds the credentials to do so. Control who can change pipeline definitions and require review. Eliminate long-lived cloud credentials in favour of federated, short-lived authentication. Scan in the pipeline — dependencies, secrets, infrastructure code, container images — and gate on the results. Protect the integrity of what is built and deployed so an attacker cannot inject a malicious artefact. And limit the pipeline’s own permissions to exactly what it needs to deploy. A compromised pipeline is one of the highest-impact positions an attacker can reach, because it has legitimate access to change everything downstream.

What they are assessing

Whether you treat the pipeline as a high-value target rather than plumbing.

What is policy as code, and why is it useful in cloud?seniortechnical

Expressing security and compliance rules as code that is automatically evaluated — in the pipeline before deployment, and continuously against the running environment. Instead of a policy document saying storage must be encrypted, you have a rule that blocks or flags any storage created without encryption. It is useful in cloud because manual policy enforcement does not scale to the speed and volume of cloud change: by the time a human reviews a configuration, ten more have deployed. Policy as code makes the rules executable, consistent and auditable, and shifts enforcement left so violations are caught before they reach production rather than found afterwards.

What they are assessing

Whether you understand automated, preventive policy enforcement.

What is software supply chain security, and why has it become prominent?midtechnical

Securing everything that goes into building and running software — third-party dependencies, build tools, pipelines, container base images — rather than only the code your team writes. It became prominent because attackers realised that compromising a widely used dependency or build system reaches everyone downstream at once, and several high-profile incidents proved the point. Modern applications are mostly assembled from other people’s code, so the majority of your attack surface is components you did not write. The practical response spans dependency scanning, verifying the provenance of what you build, generating a software bill of materials, and securing the pipeline that assembles it.

What they are assessing

Whether you grasp why supply chain became a first-order concern.

What is a landing zone, and why does it matter for cloud security?seniortechnical

A pre-configured, secure baseline environment that new workloads are deployed into — with identity, networking, logging, guardrails and account structure already set up correctly. It matters because without one, every team building in the cloud reinvents the security baseline, badly and inconsistently, and the organisation ends up with dozens of differently-misconfigured environments. A good landing zone makes the secure path the default path: teams inherit encryption, logging, network controls and policy guardrails automatically rather than having to know and implement them. It is the architectural expression of paving a safe road so people do not make their own.

What they are assessing

Whether you understand baseline-by-default as a scaling strategy for cloud security.

How would you structure cloud accounts or subscriptions for security?seniortechnical

Use account or subscription boundaries as strong isolation, because they are the clearest blast-radius boundary the cloud offers. Separate production from non-production so a mistake in development cannot reach live systems. Separate workloads or business units where isolation matters. Keep a dedicated, tightly controlled account for security and logging so that even an attacker who compromises a workload account cannot tamper with the audit trail. And manage the whole structure centrally with organisation-level guardrails that individual accounts cannot override. The principle is that the boundary between accounts is far stronger than any boundary within one, so you use it deliberately to contain compromise.

What they are assessing

Whether you use account structure as a security control rather than an accident of growth.

What logging would you enable in a cloud environment, and why?midtechnicalCommonly asked

The control-plane audit log first — the record of who did what to the environment — because it is the single most valuable source for both investigation and detection, and it is what an attacker’s activity shows up in. Then identity sign-in logs, network flow logs, and resource-level logs for sensitive services like storage and databases. Centralise them into an account or workspace an attacker cannot reach from a compromised workload, so the audit trail survives the incident it is meant to record. The common and dangerous failure is discovering during an incident that the logging needed to investigate it was never enabled, or was retained for thirty days when the intrusion began in month four.

What they are assessing

Whether you know what to log and understand log integrity and retention.

How would you detect and respond to a compromised cloud credential?midtechnicalCommonly asked

Detection leans on behaviour, because a valid credential used by an attacker is authenticated and authorised — nothing is broken, so you are looking for anomaly. Access from unusual locations or at unusual times, use of permissions the identity never exercises, enumeration activity, and creation of new access or resources are the signals. Response is faster than on-premises because you can act in the API: revoke the credential, disable the identity, and — crucially — investigate what it did using the control-plane logs, because the attacker may have created persistence such as new keys, roles or backdoor accounts. Rotating the one obvious key while leaving the roles it created behind is the classic incomplete response.

What they are assessing

Whether you understand behavioural detection and thorough response including persistence.

How would you approach a cloud migration security assessment?seniortechnical

Assess before, during and after, because the risks differ at each stage. Before: what data and workloads are moving, their sensitivity and regulatory constraints, and whether the target architecture is designed securely rather than lifted-and-shifted with on-premises assumptions intact. During: that identity, network and access are configured correctly as things land, since migration pressure is exactly when corners get cut. After: that the resulting environment matches the intended design and that nothing was left exposed in the rush — temporary access that became permanent, test data left in place, over-broad permissions granted to get things working. The recurring failure is treating migration as a networking exercise and discovering the security gaps in production.

What they are assessing

Whether you think across the migration lifecycle rather than treating security as a final check.

What is the difference between lift-and-shift and cloud-native, from a security view?midtechnical

Lift-and-shift moves existing systems into the cloud largely unchanged, which carries the on-premises security model — and its assumptions — into an environment where they may not hold. A virtual machine that assumed a trusted internal network is now in a place where that assumption is false. Cloud-native rebuilds around cloud services and can adopt cloud security properly: managed identities, platform encryption, automatic patching of managed services, fine-grained IAM. The security trade-off is speed versus fit: lift-and-shift is faster and preserves familiar controls but inherits legacy weaknesses, while cloud-native is more work but lets you build the security model the environment actually needs.

What they are assessing

Whether you understand how migration strategy shapes the achievable security posture.

A team wants to spin up cloud resources outside the approved process. How do you handle it?midscenario

Find out why first, because shadow cloud is almost always a symptom of the approved process being too slow or too restrictive rather than of people being reckless. If the sanctioned path is painful, blocking harder just drives the behaviour further underground where you have no visibility. The durable answer is making the compliant path the fast one — a landing zone with sensible guardrails so teams can self-serve safely — combined with detection for resources created outside it. In the moment, you bring the specific resources under governance and controls; structurally, you fix the friction that caused the workaround. Punishing the team without fixing the process guarantees a repeat.

What they are assessing

Whether you treat shadow IT as a process signal rather than a discipline problem.

Your cloud bill suddenly spikes. Why might that be a security concern?midscenario

Because a cost spike is sometimes the first visible sign of compromise. Cryptomining is the classic case: an attacker with valid credentials spins up expensive compute, and the bill moves before any security alert does. Data exfiltration shows up as egress charges. Resource creation in unusual regions — attackers often use regions you do not — appears as unexpected line items. So a sudden spike warrants a security look, not just a finance one: what was created, by which identity, in which region, when. Wiring cost anomaly alerting into the security triage path is a cheap and surprisingly effective detection, precisely because attackers cost money.

What they are assessing

Whether you connect financial signals to security, which many candidates miss.

How would you convince developers to care about cloud security?midcompetencyCommonly asked

By reducing the effort it takes them to be secure rather than by exhortation. Developers are measured on delivery, so security that adds friction loses to the deadline every time — the answer is to make the secure path the easy path. Provide hardened, ready-to-use templates and modules so the secure configuration is the default they inherit. Put fast, actionable feedback in the pipeline they already use, with clear fixes rather than a wall of findings. Involve them early rather than gatekeeping at the end. And frame issues in terms they own — their service, their outage, their data. Security that is somebody else’s checklist gets ignored; security built into the paved road gets used.

What they are assessing

Whether you influence through enablement rather than enforcement.

What is the biggest mistake organisations make when adopting cloud?seniorcompetencyCommonly asked

Assuming the provider’s security covers more than it does, and carrying on-premises habits into an environment that works differently. The specific failures follow from that: leaving configuration and identity unmanaged because the platform felt secure, lifting systems in without rethinking their security model, and moving faster than the governance around them. The result is the pattern behind most cloud breaches — not a sophisticated attack on the provider, but a misconfigured resource or an over-permissioned identity the customer owned and never reviewed. The organisations that do well treat cloud as a different discipline to learn rather than a data centre they happen to rent.

What they are assessing

Whether you can name the systemic adoption failure rather than a single technical one.

How do you keep current with cloud security when the platforms change constantly?entrycompetencyCommonly asked

Deliberately and selectively, because trying to track everything across multiple providers is hopeless. Follow the security-relevant announcements from the platforms you actually work with, because new services arrive with new misconfiguration opportunities and new controls. Read the provider security blogs and well-architected guidance. Understand major incidents when they happen, since cloud breaches are usually instructive about a real gap. And keep hands-on, because cloud security is learned by building and breaking rather than reading. The honest framing in an interview is that nobody is current across all of cloud, so you stay current where you work and know how to get up to speed elsewhere.

What they are assessing

Realistic currency strategy rather than a claim to know everything.

What is a cloud access security broker, and what does it do?midtechnical

A CASB sits between users and cloud services to give an organisation visibility and control over cloud usage — including the sanctioned services and the shadow ones people adopt without approval. It does four broad things: discovers what cloud services are actually being used, enforces access and data policies, protects data through controls like DLP and encryption, and detects threats such as compromised accounts or unusual data movement. Its original value was tackling shadow SaaS, and it remains relevant where an organisation needs consistent policy across many cloud services that each have their own inconsistent native controls.

What they are assessing

Whether you know the CASB category and where it fits.

What is SASE, and what problem does it solve?seniortechnical

Secure Access Service Edge combines network connectivity and security functions — secure web gateway, CASB, zero-trust network access, firewall-as-a-service — into a single cloud-delivered service. The problem it solves is that the traditional model of backhauling all traffic to a central corporate perimeter stopped making sense once applications moved to the cloud and users moved out of the office: routing a remote worker’s traffic to headquarters and back to reach a cloud app is slow and pointless. SASE moves security to the edge, close to users wherever they are, and enforces policy based on identity rather than location. It is the architectural response to the perimeter dissolving.

What they are assessing

Whether you understand the architectural shift SASE represents rather than just the acronym.

How would you secure traffic between microservices in the cloud?seniortechnical

Do not rely on network position, because a compromised service inside the network should not automatically be trusted by its neighbours. The stronger model is mutual authentication between services — each service verifies the identity of the other, typically with mutual TLS — combined with authorisation policies defining which services may call which. A service mesh can provide this consistently without every team implementing it themselves, adding encryption, identity and policy at the infrastructure layer. Network segmentation still helps as defence in depth, but the primary control is identity-based service-to-service authentication, which is zero trust applied inside the application.

What they are assessing

Whether you apply zero-trust thinking inside the environment rather than only at the edge.

What is the principle of immutable infrastructure, and how does it help security?seniortechnical

Immutable infrastructure means you never modify running servers — to change something, you build a new image and replace the old instance rather than patching it in place. It helps security several ways. Configuration drift disappears, because every instance comes from a known, scanned image rather than accumulating undocumented changes. Patching becomes a rebuild-and-replace, so systems do not fall behind. And an attacker who compromises an instance gains no persistence, because the instance is replaced regularly and any foothold goes with it. It also makes incident response cleaner: you replace rather than clean, which is faster and more reliable than trying to eradicate an attacker from a long-lived machine.

What they are assessing

Whether you understand a modern operational pattern and its security benefits.

What would you check before approving a new SaaS application for the business?midtechnical

What data it will hold and how sensitive it is, which sets the depth of everything else. Then its security posture: certifications and what they actually cover, its authentication options and whether it supports single sign-on and MFA, how it handles and where it stores data including backups, and its incident history. Then the integration: what access it wants into your environment and whether that is proportionate, since over-permissioned OAuth grants to third-party SaaS are a real and under-watched risk. Then the exit: how you get your data out and what happens on termination. The aim is a proportionate assessment matched to what the application can actually reach and hold.

What they are assessing

Whether you assess SaaS proportionately, including the integration and exit risks people forget.

How would you explain cloud security posture to a non-technical executive?seniorcompetency

In terms of exposure and trend rather than findings and services. What matters to them is how exposed the organisation is, whether that is improving, and where the biggest risks sit — not the count of misconfigurations or the names of tools. A secure-score trend over time, the number of internet-facing systems holding sensitive data, and the status of the few genuinely serious exposures communicate more than a technical dashboard. Frame it as a business would understand a risk: here is where we could be hurt, here is how likely, here is what we are doing about it, here is whether it is getting better. The failure is presenting a security console to someone who needs a decision.

What they are assessing

Translation skill applied to cloud, which senior roles require.

Identity and Access Management

35
What is the difference between identification, authentication and authorisation?entrytechnicalCommonly asked

Identification is the claim — who you say you are, such as a username or an email address. Authentication is the proof — evidence that the claim is actually true, such as a password, a token, or a biometric. Authorisation is what you are then permitted to do once your identity has been proven — which files, which systems, which actions. They are sequential and genuinely distinct, and conflating them causes real, common design errors worth being able to name: a system that authenticates strongly but authorises poorly lets the right people do the wrong things (you’ve proven who they are, then let them touch everything), while a system that authorises before properly authenticating is making access decisions on an unverified claim, which is how privilege-escalation and access-control bugs happen. A concrete way to hold the three apart: identification is you saying ‘I’m Sam,’ authentication is you proving it with your password and MFA, and authorisation is the system then deciding that Sam is allowed to read the finance folder but not the HR one. The detail candidates most often forget, and the one that impresses when included, is that there’s a fourth ‘A’ completing the set: accounting (or auditing) — the record of what was actually done once authenticated and authorised. Without it you can control access perfectly and still have no idea what happened, which is why the full model is often called AAA, and why logging is not an afterthought but the fourth pillar of the same structure.

What they are assessing

Precision on the foundational sequence, and whether you remember accounting.

What are the authentication factors, and why are they not equal?entrytechnicalCommonly asked

Three categories: something you know (a password), something you have (a token or phone), and something you are (a biometric). Multi-factor means combining categories, so a stolen password alone is not enough. They are not equal because the something-you-have factors vary enormously in strength: an SMS code can be intercepted or SIM-swapped, an app code or push is stronger but still phishable, and a hardware security key is phishing-resistant because it binds the authentication to the legitimate site so a proxy phishing page cannot relay it. The modern point is that having MFA is a solved conversation; which kind is where the current attacks are decided.

What they are assessing

Whether you understand the strength hierarchy, not just the three categories.

What is phishing-resistant MFA, and why does it matter now?midtechnicalCommonly asked

Authentication factors that cannot be captured and replayed by a phishing site, because the authentication is cryptographically bound to the legitimate domain. FIDO2 security keys and passkeys are the main examples: the key checks it is talking to the real site before responding, so a proxy page sitting between the user and the service gets nothing useful. It matters now because attackers routinely defeat app-code and push MFA with real-time proxy phishing — the user enters their code into the fake site and the attacker relays it instantly. Phishing-resistant factors are the answer to that specific and now-common attack, which is why regulators and NCSC increasingly push for them.

What they are assessing

Whether you understand the current attack that ordinary MFA no longer stops.

What is the principle of least privilege, and why is it hard to maintain?entrytechnicalCommonly asked

Granting each identity only the access it genuinely needs to do its job, and no more — so that if any single account is compromised, the blast radius is limited to what that account could legitimately touch rather than the whole estate. It’s trivial to state and genuinely hard to maintain, and understanding *why* is the substance of the answer. Access is easy to grant and nobody’s job to remove: granting solves an immediate problem (someone’s blocked, you give them the permission, they’re unblocked), while removing solves no visible problem and carries risk (take away access and you might break something nobody wants to be responsible for breaking). So the incentives are entirely asymmetric, and the result is accumulation — people change roles and keep their old permissions on top of their new ones, ‘temporary’ access granted for a project quietly becomes permanent, and over time most identities in a mature estate hold far more access than they actually use. This is privilege creep, and it’s why the gap between access *granted* and access *used* is one of the most revealing metrics in identity security. The key insight that separates a thoughtful answer from a rote definition: maintaining least privilege is much less about getting the initial grant right and much more about the ongoing discipline of *removal* — regular access recertification, automated deprovisioning when people move or leave, and tooling that surfaces unused entitlements so they can be stripped. Least privilege is not a state you configure once; it’s a decay problem you have to keep fighting, because entropy in an access model always runs toward more access, never less.

What they are assessing

Whether you understand least privilege as an ongoing removal problem, not a one-time grant.

What is privilege creep, and how would you address it?midtechnicalCommonly asked

The gradual accumulation of access an identity no longer needs, usually from role changes where new permissions are added but old ones are never removed. Over a career an individual ends up with sedimentary layers of entitlement from every job they have held, invisible until an audit or a compromise reveals that someone in marketing can still touch finance systems. Addressing it is structural: trigger access review on role change rather than relying on managers to request removal, run periodic recertification that actually surfaces anomalous access rather than dumping everything for rubber-stamping, and use tooling that highlights unused permissions. The single most effective control is recertification that people actually engage with.

What they are assessing

Whether you can name the cause and propose structural fixes rather than exhortation.

Walk me through the identity lifecycle and where it breaks.midtechnicalCommonly asked

Joiner, mover, leaver. Joiners are usually handled acceptably because someone is motivated — the person cannot work without access. Movers are where it quietly breaks: new access is added, old access rarely removed, and entitlement accumulates. Leavers are where it breaks dangerously: slow or incomplete deprovisioning leaves active credentials for people who have gone, and shared secrets they knew do not get rotated, so their knowledge outlives their account. The fix is automation driven from the authoritative source, usually HR, so that a status change triggers provisioning and, crucially, deprovisioning — rather than relying on a manager remembering to file a ticket.

What they are assessing

Whether you know which stages fail and why, not just the three names.

How would you handle deprovisioning for a departing privileged administrator?seniorscenario

Faster and more thoroughly than for an ordinary user, because a privileged leaver is a higher risk and their knowledge extends beyond their own account. Disable the account immediately on departure, ideally automatically on the HR trigger rather than on a manual request. Then the part people miss: rotate the shared and service-account credentials they had access to, because those keep working after their personal account is gone. Review what standing access and persistence they could have created — additional accounts, keys, delegated permissions. And for a contentious departure, consider doing this before they are told rather than after. The account is the easy part; the credentials they knew are the risk.

What they are assessing

Whether you go beyond disabling the account to the credentials and persistence a privileged user leaves behind.

How would you run an access recertification that people take seriously?midtechnical

Design against the reality that reviewers are busy, unmotivated, and default to approve. Make each review small and targeted rather than dumping every entitlement at once. Surface the risky and anomalous access — the person who has access nobody else in their role has — rather than treating everything equally. Give managers information they can actually interpret, in business terms, rather than raw technical entitlement names. Build in accountability so blanket approval is not frictionless and is itself recorded. The enemy is the rubber stamp, and a recertification that produces a page of approvals in thirty seconds has certified nothing.

What they are assessing

Whether you understand that the hard part of recertification is human, not technical.

What is RBAC, and when does it stop scaling?midtechnicalCommonly asked

Role-Based Access Control groups permissions into roles and assigns people to roles, so you manage a manageable number of roles rather than thousands of individual grants. It works well until two things happen. Role explosion: capturing every needed combination of access produces more roles than people, and the abstraction meant to simplify becomes its own unmanageable sprawl. And rigidity: real access needs depend on context — location, time, data sensitivity, relationship to the resource — that a static role cannot express. At that point RBAC has reached the edge of its working range, and attribute-based or policy-based models exist precisely to handle what it cannot.

What they are assessing

Whether you know RBAC as a tool with limits rather than a universal answer.

What is ABAC, and when would you use it over RBAC?midtechnical

Attribute-Based Access Control makes access decisions from attributes — of the user, the resource, the action and the environment — evaluated against policies, rather than from static role membership. You would use it where access depends on context that roles cannot express: allow access only from a managed device, only during working hours, only to records in the user’s own region, only when the data classification permits. Its strength is expressiveness and fine grain; its cost is complexity, because attribute-driven policies are harder to reason about and audit than role membership. In practice many mature systems combine the two — roles for the coarse structure, attributes for the contextual conditions.

What they are assessing

Whether you understand the trade-off and that the models combine rather than compete.

What is the difference between coarse-grained and fine-grained authorisation?seniortechnical

Coarse-grained authorisation decides access at a broad level — can this user reach this application at all. Fine-grained decides within it — can this user see this specific record, edit this particular field, approve up to this amount. The distinction matters architecturally because they often live in different places: coarse-grained authorisation frequently sits at the identity or gateway layer, while fine-grained authorisation usually has to live in the application because only the application understands its own objects and rules. A common failure is trying to push fine-grained decisions up to a layer that cannot see the context, or leaving fine-grained authorisation entirely to inconsistent per-application code.

What they are assessing

Whether you understand where different authorisation granularities belong architecturally.

What is externalised authorisation, and why is it becoming popular?seniortechnical

Moving authorisation logic out of individual applications into a dedicated policy service that applications query for decisions. Instead of each application implementing its own access rules inconsistently, they ask a central engine "can this subject do this action on this resource," and the engine evaluates policy and answers. It is becoming popular because scattered, per-application authorisation is inconsistent, hard to audit, and impossible to change centrally — externalising it gives one place to define, reason about and update policy, and one place to audit who can do what. The trade-off is a new dependency in the request path and the effort of migrating embedded logic out of applications.

What they are assessing

Awareness of a current architectural direction in authorisation.

Explain SAML, OAuth and OIDC and when each is used.midtechnicalCommonly asked

They solve related but distinct problems. SAML is an older XML-based standard for federated authentication, still very common in enterprise single sign-on to web applications. OAuth is an authorisation framework: it lets an application access resources on a user’s behalf via tokens, without handling their password — but on its own it does not authenticate the user. OIDC sits on top of OAuth and adds that missing authentication layer, which is what modern "sign in with" flows actually use. The common confusion, and a frequent security bug, is using OAuth alone for login when authentication is what is needed — OIDC exists precisely to fix that.

What they are assessing

Whether you can separate the three and know the OAuth-for-login trap.

What is a token, and why is it worth protecting?midtechnicalCommonly asked

A token is a time-limited credential representing an authenticated session or granted access — the thing an application presents to prove it may do something, so the user does not re-authenticate on every request. It is worth protecting because most tokens are bearer tokens: whoever holds one can use it, no password required. So a stolen token is stolen access, and token theft has become a primary attack because it sidesteps authentication entirely, including MFA — the attacker does not defeat the login, they steal the result of a successful one. Defences are short lifetimes, binding tokens to a device or context, and the ability to revoke them.

What they are assessing

Whether you understand tokens as the prize, and why MFA does not protect a stolen one.

What is token theft, and how would you defend against it?seniortechnical

Stealing the session token or cookie produced by a successful login and replaying it, so the attacker inherits the authenticated session without ever passing authentication themselves — which is why it defeats even strong MFA. The token is captured through a proxy phishing page, malware on the device, or a cross-site scripting flaw. Defences work on the token’s usefulness: bind it to the device or client so a stolen token fails elsewhere, keep lifetimes short so a stolen one expires quickly, and use continuous evaluation that can revoke a session mid-flight when risk signals change. Device compliance and phishing-resistant login reduce the theft opportunity in the first place.

What they are assessing

Whether you understand the post-authentication attack and layered defences against it.

What is single sign-on, and what are its failure modes?midtechnicalCommonly asked

One authentication grants access to many applications, so users authenticate once rather than repeatedly. The benefits are real — fewer passwords, better user experience, centralised control over access and deprovisioning. The failure modes are the flip side of that centralisation. The SSO identity becomes a skeleton key, so its compromise is catastrophic in a way ten separate passwords never were. The identity provider becomes a single point of failure for availability as well as security — if it is down, everything is. And any weakness in the central authentication propagates everywhere at once. Which is exactly why SSO and strong, phishing-resistant authentication are inseparable.

What they are assessing

Whether you can hold the benefit and the concentrated risk together.

What is federation, and what are its risks?seniortechnical

Federation lets identities from one domain be trusted in another, so users authenticate with their home identity provider and access resources elsewhere without a separate account — the basis of business-to-business access and much cloud single sign-on. The risk is that you are extending trust to another organisation’s identity management: if their identity provider is weak or compromised, that weakness flows into your environment through the trust relationship. You are effectively accepting their authentication decisions. So federation should come with scrutiny of the partner’s controls, scoping of what federated identities can reach, and monitoring of federated access, rather than blanket trust because a trust was configured.

What they are assessing

Whether you understand federation as extending trust and the risk that carries.

What is privileged access management, and what problem does it solve?midtechnicalCommonly asked

PAM is the set of controls specifically for privileged accounts — admins, root, service accounts with deep access — because they are a categorically different problem, not just accounts with more permissions. They are what attackers hunt, because owning one is owning the environment. PAM adds what ordinary IAM does not: vaulting and rotating privileged credentials so they are not static and shared, brokering access so admins check out elevated rights rather than holding them permanently, session recording for accountability, and just-in-time elevation so standing privilege shrinks toward zero. Framing PAM as the extra care the most dangerous accounts need, rather than "IAM for admins," is the distinction that matters.

What they are assessing

Whether you understand privileged accounts as a distinct discipline.

What is just-in-time access, and why is it better than standing privilege?midtechnical

Just-in-time access grants elevated rights only when needed, for a bounded time, with approval and logging, then removes them automatically — rather than an administrator holding privileged access permanently. It is better because standing privilege is a permanent target: an always-on admin account is exactly what an attacker wants to compromise, and its permissions are available every hour of every day whether in use or not. Just-in-time shrinks that exposure to the moments access is genuinely used, so a compromised identity most of the time holds no elevated rights at all. It also produces a clean audit trail of who elevated, when, and why, which standing access does not.

What they are assessing

Whether you understand why reducing standing privilege matters, not just what JIT is.

What is the tiered administration model, and what problem does it solve?seniortechnical

A model that separates administrative accounts and the systems they manage into tiers by sensitivity — typically domain and identity infrastructure at the top, servers in the middle, workstations at the bottom — and forbids credentials from a higher tier being used on a lower one. The problem it solves is credential theft and lateral movement: if a domain admin logs into an ordinary workstation, their credentials can be harvested there and used to take the domain. Tiering breaks that path by ensuring the most powerful credentials only ever appear on the most protected systems, so compromising a workstation cannot escalate to controlling the estate.

What they are assessing

Whether you understand the control that breaks credential-theft escalation.

What is a privileged access workstation, and why use one?seniortechnical

A hardened, dedicated device used only for administrative work, isolated from the everyday productivity environment — no email, no web browsing, no general software. The reasoning is that the workstation is where credentials are most exposed, and a normal user device is constantly under attack through email and the web, so performing privileged work from it means the most powerful credentials sit on the most-attacked machine. A privileged access workstation removes that: admin credentials are only ever entered on a device specifically hardened and kept away from the routes malware arrives by. It is a core part of protecting tier-0 access and pairs naturally with tiered administration.

What they are assessing

Whether you understand why administrative work needs a separate, hardened device.

What are non-human accounts, and why are they an IAM blind spot?midtechnicalCommonly asked

The identities that let systems, applications and automated processes authenticate — service accounts, application identities, automation credentials, API keys. They outnumber human accounts, often heavily, and they are a blind spot because IAM is built around people. They do not fit the joiner-mover-leaver lifecycle, so nothing reviews them. They are frequently over-permissioned, because broad access made an integration work and tightening it risked breaking production. Their credentials are static and rarely rotated, because rotation might break the dependent system. And ownership is murky, so when their creator leaves they become orphaned, powerful and unwatched — which is exactly what an attacker wants to find.

What they are assessing

Whether you understand why non-human identities are systematically neglected.

What is machine identity management, and why is it growing?seniortechnical

Managing the identities and credentials of non-human entities — services, workloads, devices, and the certificates and keys they use — at scale. It is growing because the number of machine identities is exploding relative to human ones: microservices, containers, serverless functions, automation and cloud workloads each need to authenticate, and they now vastly outnumber people. Managing this by hand does not scale, and the failure modes — expired certificates causing outages, leaked keys granting standing access, orphaned workload identities — are exactly the ones that cause incidents. The direction of travel is short-lived, automatically issued and rotated credentials, and workload identity that removes stored secrets entirely.

What they are assessing

Awareness of a fast-growing area and where it is heading.

What is certificate lifecycle management, and why does it cause outages?midtechnical

Managing digital certificates across their whole life — issuance, deployment, renewal and revocation. It causes outages because certificates expire, and an expired certificate on a production system stops it working: connections fail, integrations break, services go down. The reason this keeps happening is that most organisations do not actually know every certificate they hold or when each one expires, so renewal depends on someone remembering rather than on a process, and the failure is discovered at nine in the morning by the service desk rather than in advance by an inventory. The fix is discovery and automation — know every certificate and renew it automatically so expiry stops being a human memory problem.

What they are assessing

Whether you understand a mundane but high-impact operational reality.

What is conditional access, and how would you deploy it safely?midtechnicalCommonly asked

Conditional access makes authentication decisions from signals and policy rather than a static allow: who the user is and their risk level, whether the device is managed and compliant, location, and the sensitivity of what they are reaching — producing outcomes like grant, grant-with-MFA, require-compliant-device, or block. It is how zero trust is implemented at the identity layer. Deploying it safely is the real test, because one misconfigured blocking policy can lock out an entire organisation in a single click. So you never enable a blocking policy straight to enforce: report-only mode first to see what it would do, exclude break-glass emergency accounts before anything else, then pilot, then expand in stages.

What they are assessing

Whether you know both the mechanism and the safe deployment discipline.

What is a break-glass account, and how should it be managed?midtechnical

An emergency access account that stays usable when normal access mechanisms fail — when conditional access misfires, the identity provider has an outage, or a misconfiguration locks everyone out. It exists precisely so that a mistake or an incident cannot leave you with no way in. Because it is powerful and excluded from the controls that protect everything else, it needs its own careful handling: excluded from conditional access blocking so it survives a lockout, a very strong credential stored securely offline, tightly restricted knowledge of it, and alerting on any use so that legitimate emergency use is visible and misuse is caught immediately. Most organisations keep two, in case one fails.

What they are assessing

Whether you understand the account that makes aggressive access policies safe to deploy.

What is zero trust, and what does it mean for identity specifically?midtechnicalCommonly asked

Zero trust is the principle that no access is granted because of where a request comes from — being on the network stops being a credential, and every access is verified explicitly. For identity specifically, it means identity becomes the primary control plane: every request is authenticated strongly, evaluated against device health and context, granted least privilege, and re-evaluated continuously rather than trusted for the life of a session. Conditional access, phishing-resistant MFA, device compliance and continuous evaluation are the mechanisms that implement it at the identity layer. In practice zero trust is less a product than a programme of replacing every implicit trust — the flat network, the legacy protocol, the standing privilege — with an explicit, verified one.

What they are assessing

Whether you can connect zero trust to concrete identity mechanisms rather than reciting the slogan.

What is the difference between Active Directory and Entra ID?midtechnicalCommonly asked

Active Directory is the on-premises directory and authentication service that has run enterprise Windows environments for decades, using protocols like Kerberos and LDAP and organised around domains, forests and group policy. Entra ID (formerly Azure AD) is Microsoft’s cloud identity service, built for cloud and SaaS authentication using modern protocols like OIDC and SAML. They are not the same product renamed — they solve different problems and are frequently run together in a hybrid model, synchronised so users have one identity across both. Knowing they are distinct, and how they connect, matters because most UK enterprises run exactly that hybrid arrangement rather than one or the other.

What they are assessing

Whether you understand the two are different systems commonly run together, not a rename.

What is hybrid identity, and what are its security risks?seniortechnical

Running on-premises Active Directory and cloud Entra ID together, synchronised so a user has one identity spanning both. It is the dominant enterprise model because organisations have decades of on-premises investment and are adopting cloud simultaneously. The security risk is that it joins two attack surfaces and creates a bridge between them: a compromise on-premises can flow to the cloud and vice versa, the synchronisation infrastructure itself becomes a high-value target, and misconfiguration in the trust between them can let an attacker move across. The connection components — the sync servers and their accounts — are effectively tier-0 and are frequently under-protected relative to their power.

What they are assessing

Whether you understand hybrid as joining two attack surfaces rather than adding one.

Why is Active Directory such a common attack target?midtechnicalCommonly asked

Because in most enterprises it controls authentication and authorisation for nearly everything, so compromising it compromises the estate — and because it is old, complex, and full of accumulated misconfiguration nobody has cleaned up. Delegation settings, service accounts with excessive rights and weak passwords, permissions granted years ago and never reviewed, trust relationships between domains: each is a potential path, and attackers have a mature toolkit for finding and chaining them. The effort-to-reward ratio is unmatched — one well-chosen misconfiguration can turn a single compromised workstation into domain-wide control. Understanding AD attack paths is close to core knowledge for anyone in enterprise identity security.

What they are assessing

Whether you understand why AD dominates enterprise attacks, which underpins much defensive IAM work.

A business unit insists on a shared account for a team. How do you handle it?midscenario

Understand the actual need first, because "we want a shared account" is usually a solution to a real problem — shared mailbox, shared access to a tool, coverage across a rota — and there is often a better answer than a genuinely shared credential. Shared accounts break accountability: you cannot tell who did what, cannot deprovision one person cleanly, and cannot apply MFA sensibly. So the response is to solve the underlying need properly — a shared mailbox with individual access, a group granting individual identities the same rights, delegated access — rather than either flatly refusing or waving through a password everyone knows. Where a true shared account is genuinely unavoidable, it goes in a vault with checkout and logging.

What they are assessing

Whether you solve the underlying need rather than refusing or capitulating.

Users are complaining that security controls make login too painful. How do you respond?midcompetencyCommonly asked

Take it seriously rather than dismissing it, because friction that users hate gets worked around, and a control people route around protects nothing. Find out specifically what is painful — often it is frequent re-authentication or MFA prompts that could be tuned without reducing security. Then use risk-based approaches to reduce friction where risk is low: conditional access can skip prompts for a compliant device in a trusted context and demand more only when signals warrant it, so security tightens exactly where it matters and eases where it does not. The goal is making the secure path the low-friction path, because usability and security are allies more often than the framing suggests.

What they are assessing

Whether you treat friction as a real risk rather than a user complaint to ignore.

How would you approach an identity governance programme from scratch?seniortechnical

Start with visibility, because you cannot govern what you cannot see: who has access to what, including the non-human accounts, and where the authoritative source of identity is. Then the lifecycle, automated from that source so joiners, movers and leavers are handled reliably rather than by tickets. Then access reviews that actually work, targeting risky and anomalous access rather than rubber-stamping everything. Then privileged access, because that is where the concentrated risk lives. Sequence by risk and deliver visible improvements early, because governance programmes that arrive with a grand framework and no quick wins lose the credibility they need. The unglamorous foundation — knowing what exists — is where most programmes should start and often do not.

What they are assessing

Whether you can structure a programme by risk with the right foundation.

How would you measure whether identity security is actually working?seniorcompetency

By outcomes rather than deployment. Meaningful measures: the proportion of privileged access that is just-in-time rather than standing, MFA coverage across accounts including the phishing-resistant kind, the time between a leaver departing and their access being fully removed, the number of orphaned and unused accounts, and how much of the estate’s access has actually been reviewed and by someone who engaged with it. Deployment metrics — "we have MFA," "we have PAM" — measure that a control exists, not that it works. The honest measure asks whether standing privilege is actually shrinking, whether leavers actually lose access promptly, and whether reviews actually change anything.

What they are assessing

Metric literacy applied to identity, and resistance to existence-not-effect measures.

What do you see as the biggest shift happening in identity right now?midcompetencyCommonly asked

The move toward eliminating passwords and standing credentials entirely — passwordless and phishing-resistant authentication for people, and short-lived, secretless workload identity for machines. It is driven by the fact that the credential is the thing attackers steal, so the strongest defence is not to have one sitting around to be stolen: passkeys remove the phishable password, and workload identity federation removes the static key. Alongside it, the shift from perimeter to identity as the control plane, and from point-in-time authentication to continuous evaluation. The honest framing is that these are directions of travel most organisations are partway along rather than finished states, and the interesting work is the migration.

What they are assessing

Whether you have a current, defensible view of where the field is going.

Security Engineering

32
What does hardening actually mean, and how would you approach it?entrytechnicalCommonly asked

Reducing a system’s attack surface and strengthening its configuration so there is less to attack and less that will succeed. Approach it by starting from a recognised baseline rather than inventing one — a benchmark tells you what good looks like for that platform — then tailoring it to what the system actually does, because a hardening guide applied blindly breaks things. The core moves recur across platforms: remove what is not needed, close what is not used, apply least privilege, enable logging, and keep it patched. The discipline that separates engineers is doing this as a repeatable, tested baseline rather than a one-off manual effort per machine.

What they are assessing

Whether you approach hardening from a baseline and tailor it, rather than reciting settings.

How would you harden a build image so that hardening does not drift?midtechnical

Bake the hardening into a golden image that every system is built from, rather than hardening each machine after deployment. That makes the secure configuration the starting point rather than a task someone might skip. Then defend against drift two ways: scan running systems continuously against the baseline so deviations are caught, and prefer immutable infrastructure where you replace rather than patch, so machines never live long enough to drift. Build the image as code so it is versioned, reviewed and reproducible. The failure mode you are designing against is the estate where every machine was hardened once, at build, and has quietly diverged ever since.

What they are assessing

Whether you understand drift as the real enemy and design against it.

What is the principle of secure defaults, and why does it matter?entrytechnical

That systems should be safe in their out-of-the-box state, so security does not depend on every operator knowing to turn it on. It matters because most misconfiguration is omission rather than error — the port left open, the encryption not enabled, the default password not changed — and defaults are what most systems actually run with, because most people never change them. Designing for secure defaults means the safe configuration is the one you get for free and insecurity requires a deliberate choice, which inverts the usual situation where security is the extra effort. For an engineer building platforms for others, secure defaults are how you protect people who will never read the hardening guide.

What they are assessing

Whether you understand that defaults are the real configuration most of the time.

How would you approach patch management as an engineering problem?midtechnicalCommonly asked

As a reliable process rather than a tool, because everyone owns a patching tool and almost nobody has a working process. The parts that actually determine success: an accurate inventory, because you can only patch what you know exists and the unpatched machines are always the ones nobody knew about; prioritisation, so critical and exploited vulnerabilities on exposed systems jump the queue rather than everything moving at the speed of the trivial; testing proportionate to risk; deployment in rings so a small group absorbs surprises first; and verification that patches actually applied, because "we deployed it" and "it applied everywhere" are different claims. The gap between them is where compromised machines live.

What they are assessing

Whether you understand patching as an organisational process with known failure points.

How would you handle a vulnerability you cannot patch?seniortechnicalCommonly asked

Work the compensating-controls ladder in order of value. First understand why it cannot be patched, because the reason shapes the plan — vendor gone, certified against a fixed version, or just feared downtime, which is not really "cannot." Then reduce exposure: segment it hard, strip its network paths to the minimum, take it off any internet-facing route. Then restrict interaction and virtual-patch with an IPS or WAF carrying signatures for the known issue. Then monitor it disproportionately, because a known permanent weakness earns the best detection in the estate. Then formalise: a risk acceptance with an owner and review date, and a replacement roadmap. The failure mode is the temporary mitigation that quietly becomes permanent, unowned architecture.

What they are assessing

Whether you can work the compensating-controls ladder rather than shrugging.

What is virtual patching, and when is it appropriate?midtechnical

Blocking exploitation of a known vulnerability in transit — typically with an IPS or web application firewall carrying a signature for the specific issue — rather than fixing the underlying flaw. It is appropriate as a stopgap when you cannot patch immediately: a legacy system that cannot take the update, a window before a patch can be tested and deployed, or a system that genuinely cannot be modified. The honest framing, which matters in an interview, is that it is a mitigation and not a fix — it reduces the risk of the known attack path without removing the vulnerability, so it should be labelled as such and paired with a plan to actually remediate, not treated as the job done.

What they are assessing

Whether you present virtual patching honestly as mitigation rather than a fix.

Explain symmetric versus asymmetric encryption and where each is used.entrytechnicalCommonly asked

Symmetric encryption uses one shared key for both encryption and decryption — fast, but both parties need the same secret, which creates a distribution problem. Asymmetric uses a key pair, public and private, where anything encrypted with one is decrypted only with the other — slower, but it solves key distribution because you can share the public key freely. In practice they combine: asymmetric cryptography is used to establish trust and exchange a key securely, then the actual bulk data is encrypted symmetrically because it is far faster. TLS is the everyday example — asymmetric to authenticate and agree keys during the handshake, symmetric for the session that follows.

What they are assessing

Whether you understand why both exist and how they are used together.

Explain hashing versus encryption versus encoding.entrytechnicalCommonly asked

Three transformations for three purposes, and conflating them causes real insecurity. Encoding changes representation for compatibility and reverses trivially — base64 has no key and provides zero confidentiality. Encryption transforms data so it is unreadable without a key and reversible with it — confidentiality for data you need back. Hashing is one-way: any input to a fixed-size digest with no path back, used for integrity and verification. The use that matters most is passwords, which you hash — never encrypt, never encode — because the system only needs to verify a match, not recover the password, and properly means slow, salted, purpose-built password hashing. Describing base64 as encryption is the classic disqualifier.

What they are assessing

Precision on a distinction that candidates botch and that has real security consequences.

How would you store passwords securely?midtechnicalCommonly asked

Hash them with a slow, salted, purpose-built password hashing function — Argon2, bcrypt or scrypt — never encrypt them and never use a fast general-purpose hash. The reasoning: the system never needs the password back, only to verify a match, so hashing is correct and encryption is wrong because encryption implies a key that, once stolen, exposes every password. Salting with a unique per-password salt ensures identical passwords produce different hashes and defeats precomputed attacks. The slowness is deliberate — it makes brute-forcing expensive. And never roll your own: use the platform’s vetted implementation, because password hashing is exactly the kind of thing that is subtly easy to get wrong.

What they are assessing

Whether you know the correct approach and can justify each choice.

How does TLS actually work, and what breaks in practice?midtechnicalCommonly asked

TLS does two jobs. It establishes an encrypted channel — a handshake where client and server agree parameters and establish keys using asymmetric cryptography, after which the session runs on fast symmetric encryption. And it authenticates the endpoint: the server presents a certificate binding its identity to a public key, signed by a trusted authority, and the client validates it. That validation is what makes encryption meaningful, because an encrypted channel to an attacker is worse than useless. What breaks is almost always the certificate layer in unglamorous ways — expiry outages, forgotten certificates on forgotten systems, weak or sprawling wildcard use, private keys handled carelessly, and users trained to click through warnings until a real one gets clicked through too.

What they are assessing

Whether you understand both jobs of TLS and that the failures are operational.

What is a common mistake developers make with cryptography?midtechnical

Building it themselves. The recurring, expensive mistakes are almost all variants of not using vetted implementations correctly: rolling a custom scheme, using a broken or outdated algorithm, using encryption where hashing is needed or vice versa, reusing initialisation vectors or nonces, hardcoding keys in source, and getting key management wrong so the key sits next to the data it protects. Cryptographic primitives are unforgiving — a small implementation error silently destroys the security while everything appears to work. The engineering guidance is to use well-established libraries at a high level, never implement primitives yourself, and treat key management as the hard part it actually is.

What they are assessing

Whether you know that the danger is implementation, and key management is the hard part.

What is post-quantum cryptography, and should organisations care yet?seniortechnical

Cryptographic algorithms designed to resist attack by quantum computers, which would break much of the asymmetric cryptography protecting data today. Whether to care yet is the real question, and the honest answer is: not urgently for most, but not never. Large-scale quantum computers capable of this do not exist yet, but the "harvest now, decrypt later" concern is real — data with a long confidentiality lifetime, captured today, could be decrypted once the capability arrives. So organisations handling long-lived secrets should be planning crypto-agility now: knowing where they use vulnerable algorithms and being able to migrate. For most, awareness and inventory rather than immediate migration is the proportionate response.

What they are assessing

Whether you can give a measured answer rather than hype or dismissal.

What is defence in depth, with a concrete example?entrytechnicalCommonly asked

Arranging independent layers of control so that one layer’s failure is another’s catch, with no single point of failure handing over the asset. The point is that every control fails sometimes — bypassed, misconfigured, or simply missed — so you never rely on one. Concrete example, a database of customer records: it is not internet-reachable at all (segmentation); the application reaches it with a least-privileged account (so an app compromise is constrained); admin access needs MFA from a privileged workstation (so stolen credentials alone fail); the data is encrypted at rest (so a bulk read yields less); and query monitoring alerts on anomalies (so if prevention fails, you notice). Each layer fails to a different attack, which is what makes it depth rather than a pile.

What they are assessing

Whether you can make defence in depth concrete rather than reciting the phrase.

How would you segment a flat network?midtechnicalCommonly asked

As a migration problem, not a design problem — anyone can draw the target zones; the difficulty is getting an operating business from flat to segmented without breaking everything. Discovery first, because the flat network’s defining property is that nobody knows what talks to what — watch the traffic, map dependencies, over weeks. Then design zones from risk rather than the org chart. Then sequence by value, carving out the highest-stakes zone first to prove the approach. Then monitor-before-enforce: stand rules up in logging-only mode, compare what they would block against reality, chase the surprises. And plan for the residue — the legacy systems with undocumented dependencies where the project goes to die unless it is contained rather than deferred.

What they are assessing

Whether you treat segmentation as a migration with discovery, not a diagram.

What is microsegmentation, and how does it differ from traditional segmentation?seniortechnical

Traditional segmentation divides a network into zones, usually at the subnet or VLAN level, and controls traffic between them. Microsegmentation goes finer, controlling traffic down to individual workloads or applications, so that even systems in the same zone cannot talk to each other unless explicitly allowed — often enforced at the host or hypervisor rather than the network. It differs in granularity and in defeating lateral movement: traditional segmentation still leaves a compromised host free to move within its zone, while microsegmentation shrinks that to nearly nothing. The cost is complexity — far more policy to define and maintain — which is why it is usually reserved for high-value environments rather than applied everywhere.

What they are assessing

Whether you understand the granularity trade-off and its effect on lateral movement.

How would you limit lateral movement in an estate?midtechnical

Attack the things that let an intruder spread from one foothold. Segmentation and microsegmentation so a compromised host cannot freely reach its neighbours. Credential hygiene and tiered administration so a credential stolen from a workstation is not also valid on the servers — shared local admin passwords are the classic estate-wide failure, fixed with per-machine random passwords. Host-based controls and least privilege so a foothold does not become full control of the machine. And monitoring for the behaviours lateral movement produces — unusual authentication patterns, one account touching many hosts. The theme is that prevention rarely stops the initial foothold, so the containment of movement afterwards is where the real defence lives.

What they are assessing

Whether you can name concrete controls against a specific attacker behaviour.

How would you harden a new server before it goes into production?entrytechnicalCommonly asked

Start from a baseline for the platform and tailor it to the server’s role, then work the recurring moves. Remove or disable everything not needed — unused services, default accounts, sample content — because every running thing is attack surface. Apply least privilege to accounts and services. Configure the host firewall to allow only required ports. Ensure it is fully patched before exposure, not after. Enable logging and forward it somewhere the server itself cannot tamper with. Set up secure remote administration. And ideally, none of this is manual — it comes from a hardened build image so it is consistent and repeatable. The answer’s shape matters more than any single setting: a candidate reciting settings has memorised a checklist rather than understood the goal.

What they are assessing

Whether the answer has structure and reasoning rather than being a settings list.

Explain SPF, DKIM and DMARC and what each actually prevents.midtechnicalCommonly asked

Three layers that only make sense together. SPF publishes which servers may send for a domain, so a receiver can reject mail from unauthorised servers — but it checks the envelope sender, not the visible From address. DKIM cryptographically signs the message so tampering and origin can be verified — but again the signing domain need not match the visible From. DMARC ties both to the address the human actually reads, requiring alignment, and tells receivers what to do on failure. So SPF authorises servers, DKIM proves integrity, and DMARC connects them to the visible sender and adds enforcement. What none of them prevents is the lookalike domain — mail genuinely from a domain one character off passes every check, which is why authentication is necessary but not sufficient against impersonation.

What they are assessing

Whether you understand what each layer does and, crucially, what none of them stops.

How would you reduce phishing risk through engineering rather than training?midtechnical

Training helps but should not be the only line, because people will click eventually and the technical controls are what catch it. Email authentication (DMARC at enforcement) to cut impersonation of your own domain. Strong inbound filtering and link protection. Phishing-resistant MFA so that a harvested credential is not enough on its own — the single highest-value control, because it breaks the credential-phishing payoff. Attack-surface reduction on endpoints so a malicious attachment struggles to execute. Blocking or sandboxing risky attachment types. And fast, easy reporting so users become a detection source rather than only a risk. The engineering framing is that you assume the click will happen and build so it does not matter.

What they are assessing

Whether you build technical controls rather than relying on the human.

What is application allowlisting, and why is it powerful but hard?seniortechnical

Allowing only approved software to run and blocking everything else — the inverse of blocklisting, which tries to name the bad and inevitably misses new threats. It is powerful because it stops most malware by default: an executable that is not on the list simply does not run, regardless of whether anyone has seen it before. It is hard because maintaining the list in a real environment is genuinely difficult — software updates, legitimate variety, developer tools, and the constant stream of new approved applications generate exceptions, and an over-strict list blocks people’s work while an over-loose one defeats the purpose. It is most practical on systems with a stable, known software set, like servers and fixed-function machines, and hardest on developer workstations.

What they are assessing

Whether you can hold both the power and the operational difficulty honestly.

What is attack surface reduction on endpoints?midtechnical

Removing or constraining the ways malicious code can execute or spread on a device, ahead of any specific threat. Concretely: blocking Office applications from spawning child processes or executing macros from the internet, restricting script execution, controlling which USB and removable media can be used, blocking credential theft from the operating system’s credential store, and constraining the behaviours that malware relies on regardless of its specific form. The value is that it targets techniques rather than signatures, so it stops whole classes of attack rather than known samples. It pairs naturally with EDR: reduction shrinks what can happen, EDR catches what still does.

What they are assessing

Whether you understand technique-based prevention rather than signature detection.

How would you approach security automation, and what should you not automate?midtechnical

Automate the repetitive, well-understood, high-volume work where human judgement adds little — enrichment, routine remediation, evidence collection, consistency checks — because that frees scarce human attention for the things that actually need it and removes the errors that come from doing the same task manually a thousand times. What you should be cautious automating is anything where a wrong decision is costly and context-dependent: automated response that isolates production systems, destructive remediation, or anything acting on low-confidence signals. The principle is automate the decision you would make the same way every time, and keep a human in the loop where the right answer depends on judgement the automation cannot exercise.

What they are assessing

Whether you can reason about what to automate rather than automating everything.

What is infrastructure as code, and what are its security implications?midtechnicalCommonly asked

Defining infrastructure in code that is version-controlled and deployed through automation rather than configured by hand. The security upside is large: the configuration can be reviewed before deployment, scanned automatically, versioned so every change is auditable, and made consistent so the secure configuration is what deploys everywhere by default. The downside is symmetrical — a mistake in a widely used template deploys everywhere it is used, so one insecure default becomes an estate-wide exposure, and the state files and pipelines become high-value targets that can expose or alter the whole environment. Infrastructure as code makes security either much better or much worse; it rarely leaves it unchanged.

What they are assessing

Whether you see both the leverage and the multiplied risk.

How would you secure a deployment pipeline?midtechnical

Treat it as production, because it can deploy to production and holds the credentials to do so — a compromised pipeline is one of the highest-impact positions an attacker can reach. Control who can change pipeline definitions and require review. Remove long-lived deployment credentials in favour of short-lived, federated authentication. Scan within the pipeline — dependencies, secrets, infrastructure code, images — and gate on the results. Protect the integrity of build artefacts so nothing malicious can be injected. And scope the pipeline’s own permissions tightly to exactly what it needs to deploy. The mistake is treating the pipeline as trusted plumbing rather than as the powerful, credential-bearing production system it actually is.

What they are assessing

Whether you recognise the pipeline as a high-value target.

How do you evaluate a new security tool beyond the vendor demo?seniortechnical

Start from the problem, not the product: a written requirement the tool must meet, because without one the vendor’s feature list quietly becomes your requirements and you buy something impressive rather than something useful. Then a proof of value in your environment — your data, your noise, your integrations, run by the team who would operate it, for long enough to see it on a bad day rather than in a curated demo. Measure coverage against the requirement, false-positive load at your scale (because the cost is analyst hours, not the licence), integration with what you already run, and operational ownership. Then the comparisons people skip: against the tool you already have, and against not buying at all.

What they are assessing

Whether you can evaluate rigorously rather than being sold to.

What is the difference between a control that exists and a control that works?midtechnicalCommonly asked

A control that exists is an artefact — deployed, configured, on the architecture diagram, passing the audit, which largely verifies existence. A control that works produces its intended effect against a real attempt, now. The gap between them is where most serious incidents actually live, and it opens without negligence: drift, as small changes hollow the control; exceptions that never expire; decay of the surroundings, so the control still works on an estate that has moved past it; and the deployment-versus-operation gap, where the alert fires perfectly into a queue nobody reads. The only honest way to tell the difference is to test the effect — restore the backup, attempt the blocked action, trace the alert to a human acting — rather than confirming the control is present.

What they are assessing

Whether you understand that presence is not effect, and testing is the only proof.

How would you test that a control actually works?seniortechnical

Test the effect, not the existence, and there is no honest shortcut. Restore from the backup rather than inspecting the backup job. Attempt the action the control should block, from where an attacker would attempt it. Run the phish against the protected users; walk the attack path the segmentation should cut; trace one alert from firing to a human taking action. The sophistication matters less than the orientation — effect rather than artefact. And build the testing into a routine, because a control verified once is a control that existed once: the drift that opens the gap does not stop for your test, so a control confirmed a year ago is a control you are trusting on faith today.

What they are assessing

Whether you reach for effect-testing and understand it must be ongoing.

What does it mean for a control to fail safe, and why does it matter?seniortechnical

Failing safe means that when a control breaks or is uncertain, it defaults to the secure outcome rather than the permissive one — access denied rather than allowed, traffic blocked rather than passed. It matters because controls fail, and the direction of failure determines whether a malfunction is an inconvenience or a breach: an authentication system that fails open lets everyone in when it breaks, while one that fails closed locks everyone out, which is disruptive but safe. The engineering judgement is that failure is inevitable, so you design which way it falls. The tension, worth naming, is availability — fail-closed can cause outages — so the right default depends on what the control protects and what an outage costs.

What they are assessing

Whether you understand fail-safe design and the availability tension in it.

A critical patch is available but the system owner refuses the downtime. How do you handle it?midscenarioCommonly asked

Reframe it from a technical demand into a priced risk decision, because "we are afraid of the downtime" is not really "cannot patch" — it is a trade nobody has costed. Lay out both sides concretely: the cost and likelihood of the vulnerability being exploited, against the cost of the maintenance window. Often that reframing resolves it, because the owner was weighing a certain small disruption against an abstract risk and simply had not seen them side by side. If they still decline, offer interim mitigation — virtual patching, tighter exposure — and route the residual risk as a formal, owned, time-bound acceptance to whoever is entitled to make it. What you do not do is either force it unilaterally or let it drop silently.

What they are assessing

Whether you can turn a technical standoff into a governed decision.

How would you make the secure path the easy path for developers?midcompetencyCommonly asked

Remove the friction rather than adding exhortation, because security that is effortful loses to the deadline and gets routed around. Provide hardened, ready-to-use building blocks — golden images, pre-approved infrastructure modules, libraries with the safe defaults already set — so that the easy thing to reach for is also the secure one. Put fast, actionable feedback in the pipeline they already use, with the specific fix rather than a wall of findings. Automate the checks so being secure is not extra work. The paved-road idea is that you make the well-lit path so convenient that going off it is the harder choice, which achieves far more than any policy telling people to be careful.

What they are assessing

Whether you influence behaviour through enablement rather than mandate.

How do you keep your technical skills current as an engineer?entrycompetencyCommonly asked

Deliberately, and by building rather than only reading. A home lab where you can stand things up, break them, and test controls teaches more than any course. Following the platforms you actually work on, because new services and new attack techniques arrive constantly and stale knowledge shows immediately. Reproducing published techniques and understanding incidents properly — not the headline but the mechanism — so you learn why something works, which is what lets you defend against its relatives. The honest interview framing is that nobody stays current across all of security, so you go deep where you work and know how to get up to speed elsewhere, and you can point to something specific you have actually built or tested recently.

What they are assessing

Genuine, evidenced currency rather than a list of subscriptions.

What is the most over-engineered security control you have seen or can imagine, and what is wrong with that?seniorcompetency

The failure is a control so complex or burdensome that it is operated incorrectly or worked around, ending up less effective than a simpler one honestly applied. A perfect design that a busy team cannot run under pressure protects less than a cruder design they can actually sustain — because now there is a strong control everyone believes in and nobody is really operating. The lesson, which separates experienced engineers from clever ones, is to design for the operator: can a normal team run this on a bad day, will it fail safe when they do not, is the secure path also the workable one. Trading theoretical strength for real robustness is usually the right call, not a compromise.

What they are assessing

Whether you value operability over elegance, which is a mark of engineering maturity.

Security Architecture

27
What does a security architect do that a senior engineer does not?seniortechnicalCommonly asked

The unit of work changes. Engineers own components and make them work; architects own systems and make them coherent, which means the deliverables are decisions and trade-offs rather than builds. Three things mark the shift: time horizon, because the architect is accountable for how a design looks in three years rather than three sprints; breadth of currency, because an architecture decision is argued in cost, delivery risk and regulatory exposure as much as in threat; and the uncomfortable one, that the architect’s output is mostly influence, because they rarely control the teams who implement what they design. Someone who describes architecture as engineering with better diagrams has not made the shift.

What they are assessing

Whether you understand architecture as a change in the kind of work, not a seniority label.

How do you influence teams you do not control?seniorcompetencyCommonly asked

By being useful before being right. Architects rarely have authority over the teams who implement their designs, so influence is the actual mechanism, and it is earned by being in the conversation early and helping rather than gatekeeping late. Concretely: engage at design time when changing direction is cheap, frame requirements in what the team cares about rather than in security language, offer paths rather than verdicts, and build a track record of your involvement making projects better rather than slower. The architect who is known for finding routes gets invited in; the one known for saying no gets routed around, and then decisions happen without them, which is the worst outcome.

What they are assessing

Whether you understand influence as the core mechanism, earned through usefulness.

What makes a good architecture decision, and how would you record it?seniortechnical

A good architecture decision is explicit about the trade-off it makes and the reasoning behind it, because architecture is a sequence of choices under uncertainty and the reasoning matters more than the choice — it is what lets a future team understand why, and revisit it when circumstances change. Record it in a lightweight, durable form: the context, the options considered, the decision, and the consequences accepted. The discipline of writing it down forces clarity and prevents the organisational amnesia where nobody remembers why something was done and so nobody dares change it. An undocumented decision becomes folklore, and folklore is what estates calcify around.

What they are assessing

Whether you value recorded reasoning and know the lightweight mechanism for it.

Walk me through how you would threat model a new system.seniortechnicalCommonly asked

Four questions in order. What are we building — an honest picture of components, data flows and, crucially, trust boundaries, because threats live where data crosses between things that trust each other differently. What can go wrong — walking the threats systematically, using a framework like STRIDE as a completeness prompt rather than the method itself. What are we doing about it — every credible threat gets a decision: mitigate, accept explicitly, or redesign so the threat cannot exist, which is the architect’s favourite because a threat designed out never needs patching. And did we do a good job — revisit when the design changes and check against reality later, because pen test findings and incidents are the threat model’s exam results. Do it with the delivery team, not to them; they know where the bodies are buried.

What they are assessing

Whether you have a repeatable method and hold the framework lightly.

What is a trust boundary, and why does it matter in design?midtechnical

A point where data or control crosses between things that trust each other differently — internet to application, application to database, one service to another, user to administrator. It matters because threats concentrate at boundaries: within a zone of uniform trust there is comparatively little to check, but every crossing is a place where an assumption might be wrong and where validation, authentication and authorisation belong. Identifying the boundaries is often the single most productive step in a design review, because it turns a vague "is this secure" into specific questions about what is verified at each crossing. The bugs that matter usually live exactly where a boundary was assumed but not actually enforced.

What they are assessing

Whether you understand why boundaries are the focus of secure design.

How would you design security for a customer-facing web application?seniortechnicalCommonly asked

State assumptions and ask a couple of questions first, because the right design depends on context — what data, which users, what threat, what regulatory constraint. Then design in threat-informed layers. Identity first, because for a customer-facing app it is the front line: strong authentication, sound session management, defence against the credential attacks these apps actually face. The application tier: secure development underneath, with a WAF as defence in depth rather than the primary control, and secrets handled through managed identity so there is nothing stored to steal. The data layer: encryption, least-privilege database access so an app compromise is contained, and tested isolated backups. Infrastructure: segmentation so the database is not internet-reachable by design. And operations, where designs live or die: logging, monitoring, a tested incident path, patching. Name the trade-offs rather than gold-plating.

What they are assessing

Whether you can impose structure on an open design problem and name trade-offs.

What is secure by design, and how is it different from adding security later?midtechnical

Secure by design means security is a property built into a system from its inception — in the architecture, the technology choices and the data flows — rather than a layer applied afterwards. The difference is not just timing but what is achievable: some security properties can only be designed in, because bolting them on later means fighting the existing structure. A system whose trust boundaries and data segregation were designed correctly is secure cheaply; the same system retrofitted needs expensive compensating controls to approximate what the design should have given for free. The economic argument is decisive — designing security in costs a fraction of adding it after, which is why the architect’s influence at design time is worth so much.

What they are assessing

Whether you understand that some security can only be designed in, not added.

What is the shift-left principle, and where does it stop being useful?seniortechnical

Moving security activity earlier in the lifecycle — into design and development rather than testing and operations — because issues found earlier are far cheaper to fix and some can only be prevented at design time. It is sound and has genuinely improved practice. Where it stops being useful is when it is taken to mean security is entirely a build-time concern, which it is not: controls still have to operate correctly in production, detection and response only exist at runtime, and the exists-versus-works gap opens during operation, not design. So the mature view is shift-left and keep-right — do as much as possible early, but do not mistake a well-designed system for one that is being securely operated. Both ends of the lifecycle matter.

What they are assessing

Whether you can endorse shift-left while seeing its limit.

What is zero trust, beyond the marketing?seniortechnicalCommonly asked

The principle that no access is granted because of where a request comes from — being on the network stops being a credential, and every access is verified explicitly on identity, device health and context, granted least privilege, and designed assuming breach. Two corrections cut through the marketing. It is not a product: nobody sells zero trust, they sell mechanisms that implement pieces of it — strong identity, conditional access, segmentation, device compliance, monitoring — so "a zero trust solution" on a purchase order is a category error. And it is not a deployment but a retirement programme: you go through the estate finding every place trust is implicit, the flat network, the legacy protocol, the standing privilege, and replace each with an explicit verified one. Most organisations are years into that road and will be for years more.

What they are assessing

Whether you can strip the marketing and describe zero trust as a programme, not a product.

How would you approach a zero trust migration realistically?seniortechnical

As a phased retirement of implicit trust, prioritised by where that trust is most dangerous, not as a project with an end date. Start by finding the implicit trusts — the flat network segments, the systems that trust anything on the LAN, the standing privileged access, the legacy protocols that skip modern authentication — and rank them by risk. Then replace the worst first with explicit verification, prove the approach, and expand. Identity usually comes early because it is the new control plane and delivers broad benefit. Throughout, the honest status report is which implicit trusts you have retired and which remain, not a percentage on a slide. Expect it to take years, and expect some legacy trust to survive because retiring it costs more than it is worth.

What they are assessing

Whether you can sequence a large programme by risk and be honest about its duration.

Why is identity increasingly the primary control plane?midtechnical

Because the network perimeter has dissolved. Applications moved to the cloud, users moved out of the office, and devices connect from anywhere, so there is no longer a meaningful inside to defend by location — the thing that decides whether a request succeeds is who is asking and in what context, which is identity. This is why the highest-impact modern attacks are identity attacks, and why the controls doing the work that firewalls used to are identity controls: strong authentication, conditional access, device compliance, least privilege. The architectural consequence is that identity is where design attention and investment should concentrate, because it is now the boundary that actually matters.

What they are assessing

Whether you understand the perimeter-to-identity shift and its design consequence.

How do you say no to a project without becoming the department that says no?seniorcompetencyCommonly asked

By making "no" the rarest word you use and replacing it with pricing. A bare no is a dead end that teaches projects to stop asking, and a security function projects route around is worse than none because the decisions then happen invisibly. So the default answer to a risky proposal is "not that way — here are two routes that get you what you actually need, and here is what each costs." That keeps you in the conversation as the person who finds paths. When it genuinely is a no — a regulatory breach, a catastrophic risk — say it plainly, in writing, with the risk sized, and route it as a decision to whoever owns that risk. And spend hard nos like scarce currency, because an architect who says no weekly is background noise by spring.

What they are assessing

Whether you can hold the line where it matters while staying in the conversation.

How do you balance security against delivery pressure?seniortechnicalCommonly asked

Reject the framing first, because security versus delivery is mostly a false fight — both want a system that works and keeps working, and a breach is a delivery failure that arrives later with lawyers. Then be honest about which requirements are which. A small number are genuinely non-negotiable, usually regulatory or catastrophic, and you hold those firmly and say why. But most security requirements are risk trades with a time dimension, and treating all of them as hard gates is how security earns its blocker reputation and loses credibility on the few that actually are hard. For the negotiable majority, think in sequencing: ship with a basic control now and mature it, or accept a time-bound risk explicitly at the right level. And own your own friction — if your process is slow because nobody streamlined it, that is your problem to fix.

What they are assessing

Whether you can distinguish the non-negotiable few from the negotiable many.

How would you handle a business-critical system that cannot be patched?seniortechnicalCommonly asked

Work the compensating-controls ladder in order of value, after first establishing why it cannot be patched, because the reason shapes the plan. Reduce exposure first — segment it hard, strip its network paths, take it off any internet-facing route — because a vulnerability nobody can reach is a fraction of the risk. Then restrict interaction: minimum accounts, controlled administrative access. Then virtual-patch where an IPS or WAF can block the known exploit in transit, labelled honestly as mitigation. Then monitor it disproportionately, because a known permanent weakness earns the best detection in the estate. Then formalise: a risk acceptance at the right level with an owner and review date, and a replacement roadmap. The failure mode is the temporary mitigation that becomes permanent, unowned architecture.

What they are assessing

Whether you can work the full ladder and formalise the residual risk.

What frameworks do you actually use, and how?seniortechnicalCommonly asked

A few, for different jobs, used as tools rather than scripture. NIST CSF for structuring and communicating a programme, because its functions give non-specialists something to hold. ISO 27001 where a certifiable management system is needed. In the UK, NCSC guidance and the Cyber Assessment Framework in public sector and critical services. What matters more than the list is the how: frameworks are completeness checks so you do not forget a category, shared language so "strong on protect, weak on detect" means the same to everyone, and a credibility shortcut with auditors and boards. What you avoid is implementing one literally as a substitute for thinking, because a framework describes what good looks like in general, not how to get there in your specific estate — and an organisation that implements every control without judgement is compliant and still breachable.

What they are assessing

Whether you use frameworks for orientation rather than as scripture, and know the UK ones.

How would you assess the security architecture of an acquisition?seniortechnical

Two different jobs with different clocks. Before the deal closes, due diligence asks not "is their security perfect" — nobody’s is — but "is there anything here that kills the deal or changes the price": an active or unhandled breach, because you can buy someone else’s incident; major compliance exposure; systemic neglect suggesting the estate is worse than the diagram. With limited time and access, prioritise ruthlessly — crown jewels, known incidents, identity, internet-facing exposure. After the deal, integration is where the most-missed risk lives: the moment you connect their environment to yours, their weaknesses become yours. So connection is a deliberate, staged decision, not a default — assess before connecting, keep them segmented until they meet a baseline, remediate as a funded programme. The classic failure is connecting fast to realise synergies and inheriting an intrusion you never assessed.

What they are assessing

Whether you separate due-diligence from integration and name the integration risk.

How would you design for resilience rather than just prevention?seniortechnical

Start from the assumption that prevention will sometimes fail, because it will, and design so that failure is survivable rather than catastrophic. That means containment by design so a single compromise cannot spread — segmentation, blast-radius limits, least privilege between components. It means detection and response as first-class design concerns, not afterthoughts, because if prevention fails the question becomes how fast you notice and contain. It means tested recovery — isolated backups you have actually restored, not just configured — because the ability to recover is the final control when everything else has failed. And it means graceful degradation, so a system under attack loses function rather than falling over entirely. Resilience is the acknowledgement that security is not only about keeping attackers out.

What they are assessing

Whether you design assuming breach rather than only trying to prevent it.

How would you limit the blast radius of a compromise?seniortechnical

Design so that compromising one thing does not compromise everything, which is mostly about segmentation and privilege. Network and workload segmentation so a foothold cannot freely reach the rest of the estate. Least privilege everywhere, so a compromised identity or service holds only what it needs and a stolen credential is constrained. Strong boundaries between environments and tiers, so a development compromise cannot reach production and a workstation compromise cannot reach domain control. Isolation of the most valuable assets behind additional controls. The mental model is that you cannot prevent every compromise, so you design the estate as a set of compartments — like bulkheads in a ship — so that flooding one does not sink the whole.

What they are assessing

Whether you think in compartmentalisation as a design discipline.

How would you prioritise across a multi-year security roadmap?seniortechnical

By risk reduction per unit of effort, honestly assessed, rather than by what is interesting or what a framework lists. Start from the actual threat picture and the crown jewels, and prioritise the work that most reduces the likelihood or impact of the things that would genuinely hurt. Favour foundational work that many other improvements depend on — identity, visibility, asset inventory — because building on a weak foundation wastes the later effort. Sequence for early credibility too, because a roadmap that delivers nothing visible for eighteen months loses the support it needs. And stay adjustable, because the threat landscape and the business both change, and a roadmap defended past its usefulness is its own risk. The discipline is resisting both the shiny and the merely compliant.

What they are assessing

Whether you prioritise by risk-reduction-per-effort with foundations first.

How would you assess and communicate technical debt in security terms?seniortechnical

Frame it as accumulated risk with an interest rate, because that is what it is and it is language a business understands. Security technical debt — the unpatched legacy system, the flat network never segmented, the authentication never modernised — is risk that compounds: it gets more dangerous and more expensive to fix over time, and it constrains everything built on top of it. Assess it by its exposure and by what it blocks, not just its age. Communicate it by making the ongoing cost visible: not "this system is old" but "this system cannot be secured properly, increases our exposure every month, and each new thing we build around it inherits the problem." The goal is to convert invisible accumulated risk into a decision someone is accountable for.

What they are assessing

Whether you can make invisible accumulated risk into a business decision.

How do you present security risk to a board?seniorcompetencyCommonly asked

Lead with business consequence, not mechanism. A board does not need to understand the vulnerability; it needs to know what could happen, how likely, what it would cost, and what decision is being asked of it. Give them a decision rather than information: here is the exposure, here are the options with their costs, here is the recommendation. Use their vocabulary — financial impact, regulatory exposure, customer harm, operational disruption — and drop the acronyms entirely. Show trend, because direction matters more than a snapshot, and be honest about the bad numbers, because a board that senses management is always green stops trusting the report. The discipline of getting it into two minutes forces the clarity a twenty-minute technical briefing avoids.

What they are assessing

Translation skill, which is the defining senior-architect capability.

How would you explain a complex architecture decision to a non-technical stakeholder?seniorcompetency

Anchor it to something they care about and strip everything else. The stakeholder does not need the design; they need to understand the choice being made, what it costs, what it buys, and what you recommend. Use an analogy if it genuinely clarifies rather than to show cleverness, translate every technical consequence into a business one, and present it as a trade-off with a recommendation rather than a lecture. The test is whether they can make or endorse the decision afterwards — if they cannot, you explained the architecture rather than the choice. The skill is compression without distortion: keeping the decision’s real weight intact while removing everything that does not help them decide.

What they are assessing

Whether you can compress a technical decision to its business essence.

What is a security decision you would make differently now, and why?seniorcompetencyCommonly asked

The honest answer is a pattern rather than one day: earlier on, over-valuing elegance and under-valuing operability. The strong design — the tighter control, the more complete architecture — that then degrades in production, gets bypassed, or generates so much friction the team works around it, so the beautiful design protects less than a cruder one people can actually live with. What changes with experience is treating a design as done when it is correct and operable and likely to survive contact with a busy team on a bad day, rather than merely correct. You start designing for the operator: can a normal team run this under pressure, does it fail safe when they do not, is the secure path also the easy one. That trade — slightly less impressive, far more likely to still be protecting something in three years — is the whole difference between a clever architect and a useful one.

What they are assessing

Whether you can reflect on genuine growth rather than offering a humble-brag.

What is the most over-engineered security control you can imagine, and what is wrong with it?seniorcompetency

The failure is a control so complex or burdensome that it is operated incorrectly or routed around, ending up less effective than a simpler one honestly applied. A design that is perfect on paper but that a busy team cannot run under pressure protects less than a cruder one they can sustain — because now there is a strong control everyone believes in and nobody is really operating, which is worse than a modest control everyone understands. The deeper point is that the exists-versus-works gap is usually opened at design time, by architects who optimise for how good a control is rather than whether it will work in the hands it is handed to. Designing for the operator, and trading theoretical strength for real robustness, is maturity rather than compromise.

What they are assessing

Whether you value operability over elegance, the mark of a mature architect.

What is the difference between a control that exists and a control that works, from an architecture view?seniortechnical

A control that exists is on the diagram, deployed and passing the audit; a control that works produces its intended effect against a real attempt. From an architecture perspective the important insight is that the gap between them is usually opened at design time — by designs that optimise for how good a control looks rather than whether it will operate correctly in production, by controls whose success depends on effort nobody will sustain, and by architectures that assume operation the organisation cannot actually deliver. So the architect’s job includes designing controls that will still work after two years of drift and in the hands of a busy team, and designing in the verification — the testability — that lets someone confirm they still work rather than assuming it.

What they are assessing

Whether you connect the exists-versus-works gap to design decisions specifically.

What is the biggest weakness of compliance-driven security architecture?seniorcompetencyCommonly asked

That it optimises for demonstrable conformance rather than actual resilience, and the two diverge. A compliant architecture has evidenced that specified controls exist; it has not demonstrated that they work against a real adversary, still work after years of drift, or cover the risks the standard never anticipated. Compliance also creates a ceiling, because effort stops when the requirement is met, and it can actively distort design toward what is auditable rather than what is effective. The honest position is that compliance is a useful floor — it forces baseline hygiene and gives security leverage it would not otherwise have — but an architecture built to pass an audit rather than to withstand an attacker is the error the profession keeps repeating, and the architect’s job is to build for both while never confusing the two.

What they are assessing

Whether you can hold compliance value and its limits together at architecture level.

How do you keep an architecture practice current without chasing every trend?seniorcompetency

By distinguishing durable principles from passing fashion, and investing accordingly. The fundamentals — least privilege, defence in depth, assume-breach, secure-by-design — do not change, and a practice grounded in them ages well. What changes is the implementation and the threat landscape, so you track those selectively: the shifts that genuinely alter design, like identity becoming the control plane or supply chain becoming a first-order concern, rather than every product category a vendor invents. The discipline is scepticism about anything sold as a paradigm shift, paired with genuine attention to the few real ones. An architect who chases every trend has no coherent practice; one who ignores all of them is designing for a threat landscape that has moved on. The judgement is telling the two apart.

What they are assessing

Whether you can separate durable principle from marketed novelty.

Go deeper
The list

Told when the bank grows?

New questions get added as the UK market moves. The list is how you hear about it — and about new books.

Never shared, never sold. One-click unsubscribe.
Or email contact@ronanblake.co.uk with ‘list’ in the subject.