What a managed SOC really changes, and the questions to ask before signing
Collecting logs solves a storage problem; detecting an attack takes a good deal more. What SOC, MDR and EPP actually cover, which metrics deserve a place in the contract, and the questions that reveal what an MSSP offer really is.
7 min readThe Corevia team
A company that has just installed a SIEM and connected thirty log sources feels it has solved its detection problem. What it has actually solved is its storage problem. Between a platform that retains events and a capability that identifies an attack in progress lie engineering work, a round-the-clock organisation and a mandate to act. That is what you buy when you outsource a SOC, and what the cheapest offers leave out.
What separates collection from detection
The first metric highlighted in commercial proposals is often ingested volume, in gigabytes per day or events per second. That figure mostly serves to produce the invoice. The useful question is coverage: what proportion of your estate actually emits usable telemetry, and which sources are missing? A SOC without directory authentication logs, without endpoint telemetry, without audit logs from software-as-a-service applications and without visibility on the backup console is blind in the very places where attacks unfold.
Then comes detection content, which is a discipline in its own right. A rule is designed from documented attacker behaviour (the MITRE ATT&CK matrix provides a public reference for describing it), then tested, tuned and retired once it no longer serves a purpose. A rule never run against a simulated attack remains a hypothesis. Serious SOCs validate their content through controlled tests that replay specific techniques in the client’s environment and measure what fired, what went unnoticed and what only reacted too late. When comparing two providers, that largely invisible work weighs far more than the number of rules delivered.
EPP, EDR, SIEM, SOC, MDR: who does what
EPP protects the endpoint: signatures, behavioural analysis, blocking, device and application control. It prevents, but tells you almost nothing.
EDR continuously records endpoint activity, which enables retrospective investigation and remote actions: isolating a machine, killing a process, collecting artefacts. XDR extends that correlation to identity, email, network and cloud.
SIEM centralises, normalises, correlates and retains logs. It remains a tool, and someone has to operate it.
A SOC is the organisation that operates those tools: analysts, escalation procedures, shift rotations, genuine round-the-clock cover.
MDR adds a commitment to an outcome: detecting, investigating and responding on the client’s behalf, with an explicit mandate to act on their information system.
The decisive boundary lies between SOC and MDR, and it comes down to one question: at three in the morning, when an executive’s laptop shows signs of active encryption and nobody answers the phone, does the analyst isolate the machine or simply open a ticket? Both answers are defensible, but they do not cost the same and do not carry the same liability. They belong in the contract as a list of pre-authorised actions, the people entitled to request them, and the procedure that applies when nobody can be reached.
The metrics that matter
Mean time to detect, MTTD, only means something once its starting point is defined. Measured from alert creation, it assesses the speed of a console. Measured from the attacker’s first observable action in the telemetry, as reconstructed after the fact during the investigation, it assesses the defence. The second definition is the one to write into the contract, accepting that it is less flattering: it includes the days an intrusion went unnoticed for want of a collected source or a suitable rule.
Mean time to respond suffers from the same vagueness, made worse by the ambiguity of the “R”: responding, containing, eradicating or restoring service are neither the same operations nor the same durations. A provider committing to a thirty-minute MTTR is usually committing to alert acknowledgement, which is useful but very different from effective containment. Define each milestone separately: acknowledgement, qualification, containment, closure.
False positive rate and precision: the share of escalated alerts confirmed as real incidents. A high rate exhausts teams and eventually produces real incidents that get ignored.
Telemetry coverage: the share of inventoried assets actually emitting, source by source. The inventory is the denominator; without it the figure means nothing.
Detection coverage for the techniques relevant to your context, leaving aside the rest of the matrix, whose full coverage has no operational meaning.
Share of incidents discovered by the SOC as opposed to those reported by a user, a client or a third party: the hardest metric of all to dress up.
Time to onboard a new source, and time until the first tuned detection on that source.
Alerts handled per analyst per shift: beyond a certain point, qualification becomes mechanical.
A monthly report announcing millions of events analysed and a few thousand threats blocked says nothing about your security posture: it describes internet background noise. A useful report lists qualified incidents and how they were handled, rules added, changed or retired during the month, sources lost or gone silent, coverage gaps identified, and recommendations that have stayed open too long. Ask to see a real, anonymised example before signing.
The questions to ask before signing
1Is response a mandate or an opinion? Which actions are pre-authorised, by whom, and what happens if nobody can be reached?
2Who covers nights and weekends: the same team, another site, an on-call rota? How many analysts, and how is the shift handover done?
3Where are my logs stored, under which jurisdiction, how long in immediate access and how long in archive?
4At the end of the contract, can I retrieve my logs and the detection rules developed for my environment, in a usable format?
5What exactly does the price cover: number of sources, volume ceiling, overage billing, number of investigations, forensic hours included?
6What does the service commitment apply to: acknowledgement, qualification or containment? Who assigns the severity level, and can I challenge it?
7Who performs the tuning during the first ninety days, is it billed, and what criterion declares the service genuinely in production?
8In a major incident: who leads, which named contacts, what surge capacity, and is in-depth forensic work included or subject to a separate order?
The first ninety days are the real project. That is when sources are onboarded, when the noise specific to your environment is filtered out, legitimate exceptions included, and when the provider learns what normal looks like for you. A service that goes live in a week with high alert volumes has been plugged in without being tuned. Plan for that period, name the person who will spend hours on it from your side, and judge detection quality only once the phase is over.