Cybersecurity Benchmarking: Why, Why Not, When and How
- Phil Venables
- 6 minutes ago
- 7 min read
tl;dr
Benchmarking is a waste of time when focused solely on inputs (e.g. budgets) rather than outcomes (e.g. effectiveness of controls). The budget comparisons are never “apples for apples” and may often end up setting risk tolerance only marginally ahead of others who may be in a bad state to begin with.
Instead, we need to decouple this and compare leading not lagging indicators of performance to show (i) how those leading indicators drive the lagging indicators in the right direction and (ii) what the relative unit costs of those leading indicators are.
Then, we need to benchmark not against each other directly but against an idealized (but not fantastical) aggregate control set. You compare yourself to the ideal and then how you rank vs. others also compared to the ideal. This way it’s still clear if you are not in great shape even if you place ahead of your peers.
A common question from many boards and executives to their CISO is: “How do we compare to others?”
They ask this question for various reasons ranging from a genuine curiosity, a desire for improvement, or wanting to know where to set the bar to not be egregiously low (or even negligent). They sometimes also want to check that they’re not over-controlling or over-spending.
When the board or executives ask this, despite it being a perfectly reasonable question from their perspective, it can often be a major irritation to security leaders. But, at the same time, many security leaders find it immensely useful to learn from each other, perhaps using some formal benchmarking to do that.
Let’s explore what benchmarking is a potential waste of time and then look at what benchmarking approaches are useful, perhaps even immensely valuable.
Benchmarking Troublespots
Security benchmarking often relies on point-in-time assessments and highly technical KPIs (Key Performance Indicators) that fail to map to practical business outcomes. The overall problem is assuming equivalence in peer comparisons. Comparisons by default fail to account for an organization's specific operational efficiency, distinct risk appetite, or underlying infrastructure complexity. Here are some of the biggest troublespots:
Lagging Indicators: Focuses heavily on lagging measures like vulnerabilities, incidents, outages, budget, headcount, and so on. You and others might want to know if you’ve had, say, more or less incidents than a peer. But this alone doesn’t tell you the root cause you can learn from.
False Equivalencies: Assumes an “apples-to-apples” comparison among peers, ignoring differences in infrastructure architecture, technical debt, and business models. It’s very easy to fall into this trap with lagging vs. leading indicators, but it can be a trap for both. For example, you compare your higher number of SOC2 issues to a peer and feel bad. But the peer has less control depth and coverage in what their SOC2 covers, so it’s inevitable they’re going to have less issues even with equivalent rigor.
Arbitrary Budgeting: Promotes adherence to generic financial targets (e.g., spending 10% of the IT budget). This doesn’t incentivize highly efficient security programs. Conversely, in some situations you may actually need to spend more than this. The right amount to spend is, of course, dependent on your actual risk and overall situation.
Compliance Bias: Encourages a compliance-driven, checkbox mentality that satisfies audit requirements but fails to address advanced or targeted threats.
Temporal Decay: Relies frequently on point-in-time assessments that rapidly lose relevance in a dynamic threat environment.
Data Obfuscation: Depends on peer data that is often aggregated or incomplete, as organizations rarely share granular internal security information.
Risk Misalignment: Fails to account for distinct organizational risk appetites, leading to the desire for controls that may be overly restrictive or too weak for your business's actual needs.
Let’s dig more deeply into the specific problems of budget benchmarking, since it’s the one that has bugged me most over the years and is emblematic of many of the problems of poorly conceived benchmarking efforts. To be fair, there are some good attempts amid a sea of haphazard approaches, but my real problem is with the very concept of these benchmarks. So much so that I think budget benchmarking has actually held back real progress for many organizations. The main reasons, again, are that a budget is an input not an outcome and there is no agreed upon taxonomy of comparison.
Security needs to be focused on outcomes. Our goal is to mitigate security risk in the best way for our organizations while balancing the needs of customers and the commercial purpose of the enterprise. Your risk is not my risk. Your business is not my business. Your threat outlook is not mine. Just because you and I spend roughly the same doesn’t mean we will get the same result. I might have different people, different issues, different established infrastructure, etc. As an aside, if your leadership says things like “we will spend whatever it takes to make sure we have no security issues”, then when an event happens does that mean you didn’t actually spend enough despite what you said? I’ve also seen some organizations actually spend too much because of excess focus on input not outcomes which resulted in wasted product spend, conflicting projects, overlapping teams and unnecessary employee churn.
Compare based on agreed taxonomy. If you’re going to collect data to do a comparison then, of course, you need to establish consistency in those units of measure, the taxonomy of terms used to establish what is being measured and an approach for context-adjustment. Budget benchmarks or even just disclosures of what your budget is are often wildly misleading. For example, most large organizations I know could legitimately represent their security budget in a 1 to 10X range i.e. I can make it the core budget of the security team (1X) or a large percentage of the whole enterprise spend (10X) if you count systems administration activities like patching, everyone’s training, cost of product reviews, access administration. Let’s look at an analogy of a credit risk department at a bank (that oversees whether loans are too risky). How would you benchmark credit risk management budgets between banks? The cost of the team, the cost of the hedges, the computational cost of the risk calculations, the time spent by loan officers’ conforming to lending standards, the opportunity cost of too stringent credit standards? Unless you’re rigorous in the terms and scope one bank could spend a much smaller amount than another, but with good effect, and look negligent compared to another bank that’s either inefficient or loads up many other costs in their benchmark submission. This is another reason to focus on outcomes.
Interestingly, when these are done well then they usually show reductions in expenditure per unit cost of control. If you're focused on budget benchmarks for which success equates to more spend then you might, paradoxically, be disappointed with actual better outcomes. I do believe comparisons / benchmarking are a useful tool for learning, for continuous improvement and sometimes for validation. Just don’t benchmark on raw spend. Instead, focus on comparing outcomes or outcomes per unit of input, not just on input. Why? Well, is spending 90% of your IT budget on security better or worse than spending 10%? I have no idea. You might be irresponsibly pouring money down the drain in the first case or being crazily frugal in the 2nd case. We don’t know without knowing the outcomes and risk profile.
Doing Benchmarking Well
Now, there are clearly huge potential benefits in conducting benchmarking not for the purpose of external comparison (am I doing better or worse than others?) but rather for internal learning (what are the best doing that I should adopt?). Yes, it’s a subtle difference. The former gathers data to “score” while the latter might gather the same data, more or less, to spot patterns to adopt. So, what might be the best ways to get the most out of benchmarking?
Leading not lagging indicators: Focus on comparing leading indicators not just lagging indicators. For example, if you compare your vulnerability counts or vulnerability resolution SLO results with another organization what predictive power does it give you if you are worse than them? Instead if you compare leading indicators like your level of software reproducibility or adherence to an immutable infrastructure design pattern you can then see what aspects of that are a practice to focus on that affects the indicators of vulnerability resolution cadence. Benchmarking lagging indicators might make you feel good (or bad) and perhaps incidentally lead to some practice learning. Doing it for leading indicators, with some corresponding lagging indicators, will for sure help you learn.
Benchmark to an aggregate ideal: If you benchmark to a range of specific companies, say peers or near peers, in your industry it tends to become a performance league table that doesn’t inform you or your leadership what you might need to do differently. Worst, for some organizations there’s a paradox that being bottom helps you justify more funding and focus but being top raises potential questions of whether you’re over-spending. This might harm investment in areas where, despite your position in the table, your risk levels show definitive need. Instead, benchmark you and a basket of companies not to each other but to an aggregate view of strong practices. Then instead of comparing to each other, compare and rank all to the aggregate idealized benchmark comparisons. This way you know where you are vs. the ideal state, and you know how close you are to that vs. your peers. This also counters the problem in the direct comparison approach that being No. 1 in a pack of failures doesn’t make you a success because you also see the comparison to the aggregate ideal.
Framework Alignment: The aggregate benchmarking can be made up of established industry standards like NIST CSF, NIST 800-53, CIS Critical Controls, or ISO 27001. But, of course, it will be better to align to more objective control oriented frameworks than risk frameworks for all the measurement nuances we discussed earlier.
Financial Quantification and Control Effectiveness: Tie financial measures (e.g. expenditure) to leading not lagging indicators and tie the leading indicators to how they affect lagging indicators. For example, data leaks (lagging indicator) are strongly correlating to the set of data protection controls (leading indicator) and your combined cost of such controls is $X. This way there are two distinct benchmarking comparisons (1) do my controls keep incidents at an acceptable level and (2) are my controls cost-effective. It could be when you compare, adjusted for scale and complexity factors, that your control set is underperforming, or your control set is performant but 10X the cost of another organization with the same effect. A single benchmark doesn’t reveal this.
Bottom line: Benchmarking is useful for establishing baselines, it should inform strategy rather than dictate it. Cybersecurity benchmarking often satisfies executive curiosity about peer comparison but frequently misleads by relying on flawed inputs like budgets and lagging indicators that ignore unique organizational contexts. Instead budgets should focus on comparing leading indicators and how effective they are driving lagging indicators of performance. Benchmarking comparisons should focus on comparison to an ideal aggregate, then measure where you place in relation to others vs. that aggregate as opposed to each other. This encourages a focus on where you’re at vs. your absolute risk tolerances rather than someone else's. It also helps you uncover the real gems in benchmarking that is to learn from others what controls move the needle the most for the least unit cost.
Comments