Ransomware Frontline Report

9. Making Countermeasures Sustainable in Day-to-Day Operations

V2 | English | Full Report

Resilience to ransomware is not something completed by a single diagnostic assessment at fiscal year-end. Because assets, users, outsourcing partners, SaaS, vulnerabilities, organizational reorganization, and business priorities all change, countermeasures need to be embedded into normal operations. This chapter presents an implementation framework for senior engineers to evaluate and improve operations.

9.1 Maturity Is Not “Number of Products”

Measuring maturity by the number of products held — EDR, SIEM, CASB, Zero Trust, and so on — overlooks exception-driven operations and the inability to recover. Here we use the following five stages.

Stage State Typical evidence Next focus
0 Unknown Cannot explain the scope of assets, identities, or recovery Only individual knowledge and outdated documents Establish owners and an inventory.
1 Partially managed Individual products/procedures exist, but exceptions are not visible Ledgers and procedures are scattered; restoration untested Connect critical operations to the management plane.
2 Standardized Minimum standards, owners, and exception records exist MFA, patching, backup, contact network Verify with actual logs and restoration.
3 Verified Improvement cycles through exercises and measurement Restoration exercises, detection verification, issue ledger Integrate outsourcing partners, cloud, and organization-wide response.
4 Adaptive Continuously reflects changes and threat intelligence Change-linked reviews, management-level metrics Contain complexity and sustain effectiveness.

Not every organization needs to aim for Stage 4. What matters is correctly understanding one’s current position and choosing the next step appropriate to business risk.

9.2 Connecting Management Metrics and Technical Metrics

For example, even if the technical team reports a “98% patch application rate,” what management wants to know is “when will critical operations come back,” “how will customer impact be contained,” and “what significant risk remains.” This 98% is a hypothetical figure used for explanation, not a measured value from any specific survey. Establish metrics that connect the two.

Area Technical metric Metric translated into business terms
Externally exposed surface Number of unowned assets, overdue exposure exceptions Likelihood that an unknown entry point can reach critical operations
Identity MFA rate for privileged identities, count of permanent privileges, count of exceptions Likelihood of losing the management plane through a single compromise
Patching Time to remediate critical vulnerabilities, unpatched exceptions Duration for which exposed/critical assets remain subject to a known danger
Backup Job success rate, restoration-test rate Confidence that priority operations can be restored within the recovery objective
Detection Time to confirm high-severity alerts, log-gap rate Time available to judge before an attack spreads
Exercises Achievement rate for contact, restoration, and business resumption Capacity to explain and resume during a major disruption

Figures are used not for punishment, but as material for decisions. If, for example, the reason for a low patch-application rate is the risk of stopping healthcare, manufacturing, or legacy operations, a measure that raises the patch rate alone would break the business. Convert this into a decision that includes isolation, alternative systems, replacement, maintenance contracts, and risk acceptance.

9.3 Exception Management Is Central to Defense

In real environments, exceptions arise — “just this month,” “the vendor needs it,” “the equipment is old,” “we’re afraid of an outage.” Prohibiting exceptions outright causes the field to create invisible operations. Good exception management is not about justifying the exception, but about making the risk visible, giving it an expiration, and choosing a compensating measure.

Required Fields for an Exception Record

  1. Target asset, service, identity, or connection destination.
  2. Reason for the exception and its business necessity.
  3. The anticipated attack surface and affected operations.
  4. Compensating mitigations (access restriction, monitoring, isolation, time limits, etc.).
  5. Approver, start date, expiration date, and re-evaluation date.
  6. Owner and schedule for the root-cause fix (replacement, decommissioning, or redesign).

If the exception ledger is held only by the security department, it drifts out of sync with changes in the field. Link it to change management, procurement, asset management, audit, or BCP so that expirations can be detected.

9.4 The Lifecycle of Outsourcing Partners and SaaS

A security questionnaire before procurement is useful, but it alone cannot protect operations after the connection is established. Confirm the following at each stage.

Stage Minimum confirmation
Before selection Data classification, administrative privilege, regional/legal requirements, incident notification, auditability
At onboarding Least privilege, SSO/MFA, logging, separation of administrators, data-egress settings
During operation Account inventory, configuration changes, connection logs, failure notification, exercise participation
During a major incident Evidence preservation, point of contact, scope of impact, fallback operations, joint communications
At termination Confirming decommissioning of accounts, tokens, certificates, integrations, and data

The DBIR’s observation of an increase in breaches involving a third party should be taken not as “eliminate third parties,” but as the practical task of managing this lifecycle. [S03]

Figure 8: The One-to-Many Impact Created by Third-Party Connections

One-to-Many Impact of Third-Party Connections

The figure does not show that centralized administration is always dangerous. It shows that, absent a design for isolation, least privilege, time limits, operation logging, emergency shutdown, and joint response, a single compromise can spread to multiple customers or departments.

(Full figure translation deferred — see note at top of this file.)

9.5 Connecting Security to Change Management

The changes that break down ransomware resilience are not only large new-system rollouts. The accumulation of small changes — an emergency VPN exposure, a monitoring exclusion, a permanent grant of privilege, a change to backup retention, a cloud sharing setting, a temporary outsourcing-partner account — becomes dangerous.

Simply including the following questions in a change request can reduce dangerous oversights.

  • Will this become reachable from outside?
  • Will new privileges, sharing permissions, or tokens be created?
  • Do logging, monitoring, backup, or recovery procedures need to change?
  • Is this a time-limited exception? Who reverts it at the end?
  • Can this change be stopped or isolated during a major incident?

9.6 Explaining Technical Debt as Ransomware Risk

Technical debt is not a matter of “old and inconvenient to maintain.” End-of-support OSes, administrators whose knowledge is not shared with others, shared identities, undocumented networks, and untested backups simultaneously increase the probability of intrusion, the scale of lateral movement, recovery time, and the difficulty of accountability.

To management, explain not just the name of the vulnerability, but in the following form.

This asset is exposed to the outside, cannot be updated, has a shared administrator identity, and its backup restoration has not been tested. If compromised, business function X would stop, and the fallback procedure could only be sustained for Y hours. Three options — isolation, replacement, or decommissioning — are presented, with their cost and timeline.

This explanation is not meant to secure budget through fear. It is meant to let the organization explicitly accept the residual risk of not choosing an option.

9.7 An Example Annual Operating Calendar

Frequency Activity Deliverable
Daily Checking high-severity authentication, management-plane, and protection-disabled events Response record, escalation
Weekly Checking newly exposed assets, critical updates, and exception expirations Diff list, notification to owners
Monthly Review of privileged/outsourcing-partner accounts, backup failures, and log gaps Remediation tickets, summary for management
Quarterly Restoration testing for critical operations, review of connected parties Exercise results, gap against RTO/RPO
Semiannual Tabletop exercise, updating the contact network and contracts/external coordination Improvement plan, approval record
Annual Overall risk assessment, BCP update, organization-wide exercise Report to the board or equivalent, plan for the following year

9.8 To Keep the 90-Day Implementation From Failing

The key to making the earlier 90-day plan succeed is not to turn the work into “the security department’s homework.” Attach a business owner, a technical owner, a deadline, evidence of completion, and residual risk to each action.

Action Business owner Technical owner Evidence of completion Residual risk
Confirm dependencies of the most critical operation Business unit head Application owner Dependency table, RTO/RPO, sign-off party Unconfirmed outsourcing-partner dependencies
Separation of privileged identities System owner IAM lead Identity list, authentication test, exception ledger Operational burden of emergency identities
Backup restoration Business owner Infrastructure lead Restoration-exercise results, business sign-off Procedure for making up final data
Remediation of the exposed surface Service owner Network lead Reachability confirmation, configuration review Time-limited exposure exceptions

Completion should not be about being able to say “it’s done” — leave evidence that a third party or a different staff member can re-confirm. This is what supports reproducibility through staff turnover or during an emergency.