Cloudy Journey
Posts on Security, Cloud, DevOps, Citrix, VMware and others. Words and views are my own and do not reflect on my companies views. Disclaimer: some of the links on this site are affiliate links, if you click on them and make a purchase, I make a commission.
Monday, August 17, 2026
Week 5
from Vimeo / OffSec’s videos https://ift.tt/iM9Oo7R
via IFTTT
⚡ Weekly Recap: VMware Exploits, Windows 0-Day, MCP Attacks, Browser Hijacks and More
The expensive attacks are not always the clever ones.
This week had plenty of proof. Exposed services got hit, old bugs found fresh use, browser sessions became attack paths, and supply-chain problems kept spreading farther than the original compromise. A lot of it came down to access that was already there and defenses that assumed nobody would look too closely.
So, nothing magical. Just a lot of small openings turning into bigger problems. Here’s what stood out.
⚡ Threat of the Week
Suspected China APT Behind Exploitation of New VMware Flaw — A suspected China-nexus APT is assessed to be behind the exploitation of a newly patched security flaw in VMware vCenter. The attacks involve the exploitation of CVE-2026-59310 (CVSS score: 9.8), a severe directory-traversal vulnerability in the VMware vCenter server that could be weaponized by a malicious actor to execute arbitrary code. In at least one compromised instance, the attacks led to the deployment of a backdoor and. a reverse SSH binary, with the attack ultimately leading to the deployment of Babuk-derived ransomware. "Based on the case we investigated, however, we do not believe ransomware was necessarily the primary objective," QUIRSO said. "To us, its deployment looks more like a smoke screen intended to distract from the underlying intrusion and, importantly, hinder subsequent forensic analysis by encrypting evidence. We therefore see the ransomware activity in this case as potentially serving the broader intrusion rather than being its ultimate objective."
🔔 Top News
- Apple macOS Flaw Exploited to Drop Crypto Miner — A recently patched security flaw in Apple macOS has come under active exploitation in the wild to deploy a cryptocurrency miner. The vulnerability in question is CVE-2026-65400 (CVSS score: 9.8), a critical authentication issue impacting the Screen Sharing component that could allow an attacker already on the network to authenticate to the built-in remote desktop feature service without valid credentials. The shortcoming was addressed as part of an emergency update in macOS Tahoe 26.6.1, macOS Sequoia 15.7.9, and macOS Sonoma 14.8.9 earlier this month. The Netherlands National Cyber Security Center (NCSC-NL) said it received a report indicating active abuse of the vulnerability across multiple systems on which port 5900 was accessible from the internet. "In all these cases, root had gained access to the affected system and placed a Monero crypto miner," the agency said.
- Lazarus Exploits New Windows 0-Day — The North Korean threat actor known as Lazarus Group has been attributed to the zero-day exploitation of a newly patched security flaw impacting Microsoft Windows to deliver a never-before-seen backdoor targeting defense and aerospace companies across France, Germany, Brazil, and India. The activity is part of Operation Dream Job, a long-running cyber espionage and social engineering campaign orchestrated by Pyongyang-backed hackers to target professionals worldwide with fake-but-compelling job offers to steal sensitive data and install malware. The attacks have been found to exploit CVE-2026-68820 (CVSS score: 7.0), a privilege escalation flaw affecting Windows Ancillary Function Driver for WinSock ("AFD.sys") that was patched by Microsoft as part of its Patch Tuesday updates for August 2026. The attacks have been observed to deliver ForestTiger and a new backdoor called Troy.
- GeoServer Patches Critical Flaw Under Attack — GeoServer has released patches for a critical SQL injection vulnerability that can lead to remote code execution (RCE). The issue, which has yet to be assigned a CVE identifier, has been patched in versions 3.0.1, 2.28.5, and 2.27.6. Per watchTowr, the vulnerability witnessed active exploitation within hours of public disclosure and that it has seen hundreds of attempts originating from a small pool of IP addresses. GeoServer project maintainers told The Hacker News that the flaw was responsibly disclosed and was scheduled to be addressed in their regular release cycle, when details of the flaw became public knowledge last week.
- Amnesia Stealer Goes Beyond Data Theft — A newly discovered macOS stealer family called Amnesia Stealer has been found to target macOS users via ClickFix attacks. The malware, besides stealing data from 16 Chromium-based web browsers as well as other sensitive information, such as passwords, cryptocurrency wallets, Apple Notes, documents, and iCloud Keychain data, includes a streaming module that allows the attacker to interactively control the victim's web browser. One notable aspect of the stealer is its ability to copy the victim's Chromium profile, including its authentication state, and load it into a headless browser on the infected system to access the authenticated sessions. The streaming module can duplicate user profiles in Chromium-based browsers, including Google Chrome, Microsoft Edge, Vivaldi, Arc, Opera, and Brave, and establish a WebSocket channel that connects to the operator's relay and receives commands, such as navigation and mouse clicks. The remote-control component is built using the Chrome DevTools Protocol (CDP). A second WebSocket channel connects to the local headless Chromium instance. "The operator receives a live screencast of the session at around 3fps and can drive it with a full input set: keyboard, mouse, scroll, navigation and tab management," Jamf said. "In effect, the remote_stream command turns an infected host into a live, operator-driven browser running the victim's authenticated sessions, which is a materially different level of access from file collection." Amnesia Stealer is the first documented macOS malware to combine a cloned Chromium profile with CDP-based, real-time remote control to allow interactive access.
- From GhostCommit to GhostSplice — A new attack technique called GhostSplice can sidestep guardrails built around AI coding assistants and parse malicious requests that are split and hidden in a different channel, such as an MCP tool description, a tool result, and a sampling message. Each of these requests is perfectly benign on its own and processed by the assistant without refusing them. "The entire attack rests on the following fact: All three of the tool channels discussed above, together with your files and your own chat, pour into one block of the assistant's memory," ASSET Group said. "There does not exist any marking that separates the content based on its respective source. Therefore, the assistant reads it all as a single page." The attack has been described as a case of cross-channel trust fragmentation. "Due to the absence of a wall between content received from different sources (e.g., different tool channels), the attacker never needs any single one of this content to look dangerous. Instead, the idea is to embed a harmless piece in each source, and the assistant stitches them back into one instruction."
- Using Chrome DevTools Protocol for Data Theft — New research from SpecterOps detailed a post-exploitation technique that allows Chromium's CDP protocol to be enabled inside a live Google Chrome or Microsoft Edge process on Windows with an end goal to steal cookies, saved data, and authenticated browser sessions provided an attacker already has code execution permissions on the compromised host. "Cookie protections like ABE and device-bound session cookies make it harder to steal and replay session material, but they do not remove the value of an authenticated browser to adversaries," SpecterOps said. "Once CDP is enabled inside a Chromium browser, an operator can use the browser context to sidestep those replay protections, access authenticated applications, and collect saved data. Enabling CDP is a reminder that the next evolution of cookie theft may not require stealing the cookie DB and ABE key at all."
️🔥 Trending CVEs
Bugs drop weekly, and the gap between a patch and an exploit is shrinking fast. These are the heavy hitters for the week: high-severity, widely used, or already being poked at in the wild.
Check the list, patch what you have, and hit the ones marked urgent first — CVE-2026-68820 (Microsoft Windows), CVE-2026-58231 (SAP Commerce Cloud), CVE-2026-48362, CVE-2026-71398, CVE-2026-27302 (Adobe), CVE-2026-20349 (Cisco Secure Firewall Adaptive Security Appliance Software and Secure Firewall Threat Defense), CVE-2026-53413, CVE-2026-53414, CVE-2026-53415 (Zoom), CVE-2026-65400 (Apple macOS), CVE-2026-20337, CVE-2026-20338, CVE-2026-20339, CVE-2026-20345, CVE-2026-20346, CVE-2026-20347, CVE-2026-20348 (ClamAV), CVE-2026-18412 (OpenCart), CVE-2026-66147, CVE-2026-66145 (SonicWall), CVE-2026-6726, CVE-2026-6727 (Trusted Platform Module 2.0 reference implementation), CVE-2026-26035, CVE-2026-70468, CVE-2026-70465 (Fortinet), CVE-2026-65640 (WordPress), CVE-2026-65321 (PyAthena), CVE-2026-43637 (Cornac), CVE-2026-63720 (datamodel-code-generator), an SQL injection vulnerability in GeoServer, and multiple vulnerabilities in WireShark..
🎥 Cybersecurity Webinars
- How to Control the Open-Source Security Debt Created by AI Coding Tools → Learn how AI coding tools are expanding unvetted open-source use, accelerating vulnerability backlogs, and weakening existing governance. This webinar shows how to measure the resulting remediation debt, connect it to breach, audit, and productivity risks, and identify which governance models can contain it without slowing development.
- AI Can Build Exploits in Minutes. Can Your Security Team Keep Up? → AI is collapsing the time between vulnerability disclosure and attack. Advanced models can now uncover flaws, generate working exploits, and chain them into complete attack paths at machine speed. This webinar presents a practical framework for gaining the visibility, context, and response speed needed to investigate and stop threats before attackers pull ahead.
📰 Around the Cyber World
- Security Flaw in FileRun — VulnCheck disclosed details of CVE-2026-14863 (CVSS score: 8.7), a high-severity operating system command injection flaw in FileRun that could lead to remote code execution. "FileRun's thumbnail extractors build shell commands by pasting the uploaded file path into a double-quoted string and handing it to exec(), and the filename sanitizer lets $() through, so a file named $(payload).mp4 runs its payload the moment a thumbnail is generated," security researcher Valentin Lobstein said. "Any authenticated user with upload permission gets code execution; when a public file request weblink exists, so does anyone who knows its token, no account required." The issue, which affects versions up to and including 2026.2.0, has been fixed in 2026.2.1.
- ClickFix Leads to ACR stealer and GhostPipe — ThreatLocker disclosed an attempted ClickFix attack that employs embedded scripts, steganographic payload extraction, and obfuscation to deploy an advanced iteration of ACR stealer and a secondary payload dubbed GhostPipe. The ClickFix attack originated from a fake CAPTCHA prompt being served on a compromised domain, resulting in the execution of a PowerShell command that downloads an MP3 file, which is then executed using MSHTA to launch a VBScript that's responsible for running intermediate payloads designed to gather system information and extract from a remotely hosted JPG file a PowerShell script. The script serves as an in-memory module shellcode launcher to deploy ACR Stealer. The malware also contacts a C2 server to fetch secondary payloads, including a PowerShell script called GhostPipe. "This seemingly unknown script performs a proxy-based AiTM attack with the sole purpose of stealing Google logins," ThreatLocker said. "The methods used are comprehensive and inherently support relaying MFA to successfully capture credentials."
- Flaw in Citrix NetScaler — Citrix appears to have silently addressed a heap overflow vulnerability in NetScaler that can be exploited to achieve remote code execution. The issue was patched as part of updates released towards the end of June 2026. watchTowr said the vulnerability likely corresponds to CVE-2026-8452, which has been described as a memory overflow vulnerability that could lead to unpredictable or erroneous behavior and denial-of-service when the appliance is configured as a Gateway or an AAA virtual server. The vulnerability, per the threat intelligence company, can be turned into code execution to drop a PHP web shell that can survive the NetScaler packet engine being respawned, and ultimately execute commands with root privileges. Shortly after details of the flaw became public, Defused Cyber said it observed active in-the-wild exploitation efforts two calendar days later.
- Ethereum Malware Loader Goes After Portuguese-Speaking Users — Portuguese-speaking users are the target of a malware loader that uses the EtherHiding technique to dynamically locate attacker infrastructure and distribute additional payloads. "The multi-stage infection chain combines obfuscated JavaScript, Node.js, DLL side-loading, and a malicious Chromium browser extension capable of targeting Chrome and Microsoft Edge to collect cookies and web storage, capture screenshots, monitor browser activity, and receive remote commands," WatchGuard Threat Lab said. The Israel National Digital Agency (INDA), in its own analysis of EtherHiding, said the technique has been used as a delivery backend, a C2 channel, a victim database, and skimming infrastructure. In another interesting twist, attackers have been found to shift to the BNB Smart Chain testnet, essentially eliminating gas fees and making the whole operation free. "Because writes there are free, unlimited, and leave no financial trail, recent reporting has found malware command-and-control backends running entirely on a testnet," INDA said.
- Thousands of Exposed Fuel Gauges Dropped from the Internet — BitSight said it observed a dramatic drop in internet-exposed Automatic Tank Gauge (ATG) systems in the U.S., with the number declining by more than half. "From March to June, there was over a 55% drop in exposed IP addresses," it said. "Globally, exposure fell 49% from that same March peak, with the U.S. accounting for most of the decline. For ten months, from June 2025 through March 2026, the U.S. held a band of roughly 4,300 to 5,300 unique exposed IPs, averaging 4,815 across 2025. Then April fell 27.6% in a single month, May fell another 31.8%, and June continued down. By June, the exposed population was 56% below the March peak."
- Phantom Enigma Campaign Targets Brazil — An active PhantomEnigma campaign has been observed abusing compromised government infrastructure and fake police-themed documents to target banking and public-sector organizations in Brazil. The phishing messages are presented as official notices and bypass email security filters to deliver a modular Node.js backdoor that collects system data, sets up persistence, and connects to rotating C2 infrastructure. It can also execute JavaScript or deliver stealers, loaders, RMM software, and other malware. Some aspects of the activity overlap with prior reports from Positive Technologies.
- F.B.I. Agent Charged With Unauthorized Crypto Withdrawals — An F.B.I. counterintelligence agent has been charged with illicitly obtaining about $1 million worth of cryptocurrency, largely through unauthorized withdrawals from a criminal target overseas. According to The New York Times, the agent claimed he had begun taking the money in late 2024 or early 2025, and made 10 or 12 withdrawals altogether by making use of a seed phrase to the suspect's account that the F.B.I. had obtained during its investigation of the individual.
- Ukraine Dismantles Fraudulent Call Centers — Ukrainian authorities disrupted 94 fraudulent call centers during a nationwide operation that involved more than 400 searches and the seizure of thousands of computers, phones, and SIM cards. Per the Ukrainian police, the call centers were associated with schemes relating to callers impersonating bank employees, fraudulent investment services, cryptocurrency platforms, and attempts to gain remote access to victims' devices or trick them into handing over sensitive data under the guise of suspicious transactions. Some cybercriminals collected personal information about prospective victims and shared or sold those records to other operations. "Call centers were staffed by administrators, operators, and other participants in the schemes, and ready-made conversation scripts, databases of potential victims, special software, and tools for hiding and further withdrawing funds were used," the police said. During the probe, 3,336 pieces of computer equipment, 1,346 phones, over 5,200 SIM cards, 90 bank cards, access to 20 crypto wallets, and 22 cars were seized. About $2 million, €64,000, cash in hryvnias, a kilogram of bank gold in bars, and jewelry were also confiscated.
- North Carolina Man Sentenced for Cyber Extortion Scheme — Cameron Curry, 27, of Charlotte, North Carolina, was sentenced to 24 months in prison for carrying out an "extensive cyber extortion scheme" against an unnamed D.C.-based international technology company. In March 2026, Curry was convicted of six counts of transmitting or willfully causing interstate communications with the intent to extort a victim company. "Curry misused his position to access the victim company's personnel and other sensitive corporate records, which he then used to carry out the cyber extortion scheme," the U.S. Justice Department said. "Curry hatched his extortion scheme after he learned that his contract was not going to be renewed and that he would no longer be employed by the company."
- ExfilSquad's Access to Data from 13 Organizations — A new analysis of data samples published by the ExfilSquad data extortion group has confirmed "they have access to sensitive data." Fortra said the breaches were most likely limited to unauthorized access of D365 instances. "The leading theory on the initial attack vector that enabled exfiltration is misconfigured Microsoft Power Page portals that allowed for public read access," it said. "The observed leaked data formats are consistent with Microsoft Dataverse exports, suggesting unauthorized read access may have been achieved, and victims found by crawling for misconfigured Microsoft Power Portals or other enumeration techniques."
- OpenAI Rolls Out Computer History in ChatGPT — OpenAI replaced Chronicle, which builds memories from screen captures to make ChatGPT and Codex more aware of context, with Computer History. "Computer History turns your activity across apps and websites into memories and a timeline that ChatGPT and Codex can reference," OpenAI said. "You can ask natural questions about recent work, pick up where you left off, understand patterns in how you work, and turn repeated workflows into skills or automations." Computer History is off by default for ChatGPT Pro, Business, and Enterprise users in the ChatGPT desktop app on macOS. The feature is reminiscent of Microsoft's controversial Recall, which attracted scrutiny for relying on periodic screenshots to capture and index relevant information. Unlike Chronicle, which also used screenshots, Computer History relies on capturing interactions (e.g., clicks, typing, keyboard shortcuts, app switches, and context) to allow the AI to understand user workflows. "Computer History records interaction events and does not capture your screen or audio," OpenAI said. "You control which apps and websites contribute, can see and pause collection from the macOS menu bar, and can inspect or delete your history at any time." That said, it's worth emphasizing that Computer History files can contain sensitive information. "They are not encrypted by Computer History, and other programs running as your macOS user may be able to access them," OpenAI cautions. "Protect your Mac account and exclude sources you do not want included." OpenAI also warned that Computer History increases the risk of prompt injection from content in apps and websites, as the AI system can follow instructions when visiting a website containing malicious instructions. Temporary interaction event files are retained for up to 48 hours before they are deleted. However, they can be used to create memory files that can remain for extended periods of time until users explicitly delete or clear them.
- China-linked LightSpy Activity Detected in Over 13 Countries — The modular implant known as LightSpy has evolved into a broader surveillance tool that's in use in more than 13 countries and regions, including Singapore, Hong Kong, the Netherlands, Pakistan, Japan, China, Malaysia, Germany, the U.S., Thailand, Indonesia, South Korea, Austria and Turkey, Arctic Wolf Labs said. This includes previously unreported router-focused capabilities along with live router implants in Europe and Africa. "The findings expand the potential impact of the surveillance framework beyond individual devices: router access can provide visibility into every device connected to a home or office network and may persist after phones are replaced, devices are factory-reset, or operating systems are upgraded," a spokesperson for the company said. LightSpy infrastructure spans several countries, including 117 servers and 35 domains impersonating Asian electronics manufacturers and router-management services. Evidence indicates that LightSpy functions as a commercialized surveillance platform, with customer branding, billing functionality, and a demonstration environment used to market the framework to prospective buyers. The platform is assessed to be the work of a Chinese contractor after one of the operators used the LightSpy administrator's panel to place an order with KFC using their real name and office address.
- Trivy Supply Chain Attack Exposed 2,500+ Companies — A new analysis from SOCRadar has revealed that 95% of organizations impacted by the LiteLLM supply chain attack earlier this year were exposed before, coinciding with the compromise of the Trivy scanner. "Organizations did not need to install LiteLLM directly to be exposed," the cybersecurity company said. "The package could arrive through frameworks such as DSPy, MLflow, CrewAI, OpenHands, and Arize Phoenix, while its payload executed at Python startup without requiring a LiteLLM import." The incident was attributed to a threat actor known as TeamPCP.
- Massive Azure Exfiltration Campaign Exposes Millions of Enterprise Records — An active Microsoft Azure exfiltration campaign is being driven by a threat actor named "TheHatman," who has "flooded" cybercrime forums with enterprise employee databases. The data is said to have been downloaded directly from the organizations’ Azure/Entra portals using compromised credentials, although the exact intrusion vector remains unknown. The campaign impacts multiple global enterprises across IT services, hospitality, telecommunications, retail, and logistics. Some of them include McDonald's, TCS, Vodafone, HCL Technologies, Kyndryl, Gap, Hexaware, and Wyndham Hotels. "Judging by the massive size of the organizations impacted, it appears highly likely that this campaign originates from targeted exploitation of Infostealer infections rather than a systemic zero-day vulnerability in Azure," Hudson Rock said.
Conclusion
That’s the week. Some attacks needed a real exploit. Others just needed an exposed system, a stolen login, or one weak link buried in the stack.
The useful part is knowing which kind you’re dealing with before it becomes your problem. Patch what matters, close what should not be public, and keep an eye on the boring stuff. It keeps winning.
from The Hacker News https://ift.tt/i9CW6hl
via IFTTT
Make zero CVEs your new default
Now in Docker AI Governance: a single searchable record of every policy decision your agents trigger, streamed to the SIEM your security team already runs, so you can show what your agents did and what your policy stopped.
Today, Docker AI Governance now streams every policy decision in your organization into the SIEM your security team already runs, with a searchable record of all of it in Docker Cloud. You can see what your agents did, and what your policy stopped them from doing.
Enforcement is step one
Somewhere in the past year, supply-chain attacks stopped being isolated incidents. The compromises now reach the tools the industry trusts to defend itself, with Trivy and KICS among this year’s targets. Mark Lechner, Docker’s Chief Information Security Officer, called the latest wave “a permanent shift in the threat landscape”, and nothing since has argued with him. Meanwhile the volume keeps climbing. Over a quarter of production code is now AI-authored, and agents pull in dependencies at machine speed. If you run a platform team or a security program, you already know how this math feels. More code, more images, more dependencies, almost none of it written by your own engineers. And all of it becomes your responsibility the moment it ships.
None of this is news to us. Securing the software supply chain is the problem we’re here to solve, and our commitment to it is absolute. The latest round of updates widens the trusted foundation Docker is building under your supply chain, and tightens how it’s enforced. More of the software inside your images is now built and patched by Docker itself. Security coverage continues after software reaches end of life. Images get tailored to your environment without losing their guarantees. And policy enforcement now reaches every developer machine. The details are below. First, where all of this is headed.
A trusted foundation for the whole supply chain

It all starts from one principle, and Docker Hardened Images was built on it. Security that doesn’t get adopted doesn’t secure anything. The entire catalog is free for every developer, because a secure baseline shouldn’t be a premium feature. Every image is compatible with Alpine and Debian, the distributions your teams already run, and Docker builds every one of them itself, from source. Adoption is a FROM-line change, not a migration project. And every image is independently verifiable, with signed SBOMs (software bills of materials) and SLSA Build Level 3 provenance, so your auditors work from evidence instead of vendor claims.
A year in, the numbers make the case. The catalog has grown past 4,000 hardened images, plus MCP servers, Helm charts, and ELS images. It draws more than 3.5 million pulls a week, with over a million builds running regularly to keep all of it patched, and open source projects like n8n run production on DHI. The catalog grows the way it always has, driven by what customers request. But the goal was never just a catalog. The goal is one trusted foundation under your whole software supply chain, where the images you run, the packages inside them, the charts that deploy them, and the tools your agents call all carry the same provenance. Security becomes the default from day one, and it holds, without asking your teams to change how they work.
Docker is leading that charge. Here’s what that looks like in practice.
Built from source, down to every package
The hardening keeps reaching deeper into the stack. Docker Hardened System Packages take hardening below the image, to the packages inside it, across both Alpine and Debian, with every package built from upstream source, patched, and maintained by Docker in the same SLSA Build Level 3 pipeline that builds the images themselves. And the repository behind them is open to more than the catalog. DHI Enterprise customers can point apt or apk directly at Docker’s hardened package repository and bring the same packages into images they build themselves, extending the hardened supply chain beyond the images Docker ships to every image your organization builds.
The coverage keeps widening. What began with Alpine now spans Debian, with Python, the catalog’s most pulled image, among the first to ship fully hardened. The work compounds every week, and the Debian and Alpine package lists are public, so you can watch the catalog harden in real time.
If you’ve spent time chasing base-image CVEs, you know why this matters. System packages are notorious for slow fixes; a patch can sit waiting on the distribution’s next release for months or years. Docker doesn’t wait. We patch at the package level, ahead of upstream when it counts, and the fix lands in every image that uses that package, in one build wave instead of image by image. Entire businesses have been built on delivering community-distribution security updates faster than the community. With DHI, that speed is included.
The guarantees hold up under inspection, too. Packages you add through DHI customization, tailoring an image to your workloads, come from that same hardened repository, not an unverified public mirror, so they are hardened system packages in their own right and the SLA that covers the base image extends through everything you add. And because one vendor stands behind the image, the packages inside it, the CVE investigation, and the patch, your auditors get a single chain of signed provenance instead of a stack of vendor assurances.
Your distribution, meanwhile, stays your distribution. Building a hardened package ecosystem from source is a serious engineering commitment, and Docker made it twice, for Alpine and for Debian, so keeping your house standard never costs you your security posture.
Patch past end of life

Production software has a habit of outliving its maintainers. Migrations wait on budgets, dependencies, and test cycles, and CVEs don’t wait with them. That’s the problem DHI Extended Lifecycle Support (ELS) exists for. It keeps end-of-life software patched, with SBOMs and provenance maintained, for up to five more years.
ELS isn’t limited to a set catalog, either. Docker watches the end-of-life calendar and builds coverage ahead of it, and anything you don’t see, you can request. MinIO is the newest addition. Upstream archived the project in February 2026, yet in the DHI catalog it lives on, patched and hardened, and your migration runs on your schedule instead of upstream’s.
Customize at scale, manage as code
Nobody runs stock images in production. You add CA certificates, agents, and the packages your applications demand. The trouble is that in most of this market, the first change you make is where the vendor’s guarantees end, and everything after it is yours to carry. DHI customization works the other way around. You define what your images need, and Docker manages the full lifecycle of your customized images, rebuilding them through the same hardened pipeline on every upstream patch. The SBOM, the attestations, and the SLA travel with the customization instead of dying at it.
Customization operates at scale, too. Bulk customizations run through the UI, CLI, and API, with YAML configuration and GitHub Actions support, so you can tailor hundreds of repositories in one pass and let the rebuilds take care of themselves. And if your platform runs on Terraform, customization is code as well. The DHI Terraform provider mirrors and customizes hardened images with the same pull requests and reviews as the rest of your infrastructure.
The savings are real infrastructure, not a rounding error. Customers tell us they’ve shut off the CI pipelines that existed only to rebuild images, because Docker rebuilds for them. The blind redeploy cadence goes with those pipelines. You ship an update when a fix actually needs to go out, knowing exactly what changed, instead of rebuilding everything on a schedule and hoping QA catches what moved.
For organizations whose data-residency requirements keep images inside the EU, EU-hosted customizations arrive in September. Your customized images will live in Docker Hub’s EU region with the same SBOMs, attestations, and SLA as everywhere else. Residency stops being the reason your hardening program waits.
Harden beyond base images
The same standard keeps moving up the stack. The catalog now carries fully supported Helm charts, so your Kubernetes deployments start hardened too. And it carries a growing set of hardened MCP servers, because the tools your agents call deserve the same scrutiny as the images they run on.
Govern it all with Docker Scout policy
Scanning tells you what’s wrong. Policy is how you keep it from shipping. And enforcement is where most supply-chain programs quietly fail, because hardened artifacts only protect you when your teams actually use them. Developers move fast and default to what works, and the developer machine is exactly where the current wave of attacks aims.
Docker Scout policy closes that gap. It evaluates flexible, customizable policies from the CLI and inside CI, and it ships with the same policies Docker uses to verify every hardened image in the catalog. The policies are written in Rego, the industry standard, and they’re portable, so the same rules that gate a build in your CI travel with your teams to every developer machine in your organization. Gating at the registry matters, but it stops at the registry; developers can route around it all day. Policy that travels to the machine is how you hold every image you run, and every image your teams build, to the bar Docker holds itself to.
It’s an additive control. It works alongside the scanners you already run, and it’s already in the Docker subscription you have.
The foundation is already in your stack
The supply-chain problem is not going to shrink. More code is coming, agents are becoming contributors, and the patch windows regulators expect keep getting shorter. Point tools won’t carry that weight. A foundation that’s secure by default will, backed by an ecosystem that keeps it that way. That is exactly what Docker’s security portfolio delivers. Hardened content on the distributions you already run, customization that keeps its guarantees, support that outlasts upstream, and policy you control, from one vendor accountable for all of it.
And none of it asks you to adopt something new. It’s all in the Docker you already run. Your builds, tools, and pipelines stay the same. Your CVE count doesn’t.
Browse the DHI catalog and pull your first hardened image today. And if you want the full story, how all of this works together, with your questions answered live, join our live webinar in early September. We’d love to see you there. [webinar registration link]
from Docker https://ift.tt/NhprAID
via IFTTT
Unisoc VoLTE Video Call Exploit Chain Can Give Attackers Full Android Kernel Access
Security researchers at SSD Secure Disclosure have published a two-stage exploit chain that achieves full Android kernel access on devices running Unisoc modem firmware through a VoLTE video call, with no fix from the chipset maker.
The advisory, published August 17, 2026, is the second stage of a chain that began in March 2026, when SSD disclosed remote code execution in the same firmware through a malformed SIP video call. Completing the full chain requires the attacker to control a private 4G cellular network and the victim to answer the incoming video call.
"We have tried to reach out to the vendor through multiple channels (email and LinkedIn) but have not been able to receive any response," SSD Secure Disclosure said in its advisory.
The March 2026 disclosure carried the same statement. The research was carried out by an independent security researcher using the handle 0x50594d.
The privilege-escalation vulnerability is classified as CWE-1189, Improper Isolation of Shared Resources on System-on-a-Chip, and no CVE identifier has been assigned as of publication.
The flaw resides in the modem firmware shared by at least three Unisoc chipsets, among them the T606 found in the Motorola E13, the T612 found in the Realme C33, and the T7250 found in the Xiaomi Redmi A5.
Unisoc, a Shanghai-based chipmaker formerly known as Spreadtrum, supplies components to brands including Motorola, Realme, and Xiaomi for devices sold across more than 140 countries, according to the advisory.
Researchers confirmed the privilege-escalation flaw on a Motorola E13 carrying a February 2025 security patch and on a Xiaomi Redmi A5 carrying a January 2026 patch.
Running the complete chain requires a modem-level foothold from the March 2026 RCE vulnerability first, along with attacker-controlled VoLTE infrastructure and a victim who answers the incoming video call.
The researchers built their proof-of-concept environment using an open-source 4G core network, a software-defined radio for the 4G radio interface, and specialized SIM cards.
Once code is running on the modem, the privilege-escalation step works by writing a full-access configuration to the modem's ARM Memory Protection Unit through coprocessor registers, mapping the entire 32-bit physical address space as readable, writable, and executable from modem context, including the pages where the Android kernel resides.
The condition making this possible is a shared physical memory space between the modem processor and the application processor within the Unisoc SoC, with no hardware-enforced boundary preventing modem-context code from modifying kernel memory.
Researchers confirmed kernel-level code execution on a test device by observing kernel log output showing that the injected payload had run.
The August 2026 Android Security Bulletin, published before this disclosure, does not address the privilege-escalation vulnerability, and no UNISOC security bulletin covers it.
A separate UNISOC advisory from October 2025, CVE-2025-31718 (CVSS score: 7.5), describes a modem input-validation flaw on the same chipset family, though it's not clear whether it corresponds to the March 2026 SSD disclosure.
Device owners currently have no available patch or mitigation and should watch for a firmware update from their device manufacturer.
The disclosure follows independent research published in November 2025 by Kaspersky ICS CERT, which documented the same architectural condition on a different Unisoc chip, the UIS7862A, found in vehicle head units. After gaining modem code execution via a separate vulnerability, the Kaspersky team was also able to reach and modify the running Android kernel by exploiting the modem and application processor's shared physical address space.
Kaspersky described one of its lateral movement paths, involving a hidden Direct Memory Access peripheral, as a hardware-level issue not fixable through a software update. The Memory Protection Unit route used in the SSD chain is in principle addressable through a firmware change, though no such update has been committed to by UNISOC.
A coordinated Unisoc modem vulnerability uncovered by Check Point Research in 2022, CVE-2022-20210, was patched by UNISOC and distributed through the Android Security Bulletin. The two currently disclosed vulnerabilities carry no such assurance.
from The Hacker News https://ift.tt/NhpfAEY
via IFTTT
Suspected China-Nexus Actor Exploits VMware vCenter Flaw, Deploys Babuk-Derived Ransomware
Cybersecurity researchers have attributed the exploitation of a newly patched security flaw in Broadcom VMware vCenter to a suspected China-nexus advanced persistent threat (APT).
The attacks involve the exploitation of CVE-2026-59310 (CVSS score: 9.8), a severe directory-traversal vulnerability in the VMware vCenter server that could be weaponized by a malicious actor to execute arbitrary code. A fix for the flaw was released by Broadcom on July 29, 2026.
German incident response company QUIRSO assessed with moderate confidence that the exploitation campaign aimed at CVE-2026-59310 is operated by a Chinese-speaking threat actor, likely working in the UTC+08:00 time zone, which is predominantly used in Chinese-speaking regions.
"This assessment is based on the convergence of Chinese-language artifacts in attacker-created scripts, apparent reuse of research from a Chinese security publication, repeated operational use of Chinese-language tools and management software, victimology excluding mainland China, and activity patterns compatible with UTC+08:00 working hours," QUIRSO researchers Maike Orlikowski, Çağatay Yürekli, and Denis Szadkowski said.
The activity, which commenced five calendar days after public disclosure of the flaw, is estimated to have compromised 361 unique victim IP addresses across 47 countries, with most of the infections scattered across Germany (55), the U.S. (41), Turkey (38), Iran (26), and France (25).
Exploitation of CVE-2026-59309
One compromised vCenter Server Appliance analyzed by QUIRSO is said to have been targeted by both CVE-2026-59310 and CVE-2026-59309, an authentication bypass that has also witnessed active scanning efforts. Evidence shows malicious activity consistent with the exploitation of CVE-2026-59309 as early as August 1, 2026, followed by the creation of an administrative account on vCenter.
That said, no login events have been observed for the legitimate administrative account that was used to create this new account. The account creation originated from the IP address 146.59.252[.]178 and also involved vSphere discovery via the REST API on August 3 using User-Agent strings like "GoodMoodle-VCFleet/1.0," in an attempt to masquerade it as VMware-related activity.
It's worth noting that VCF Fleet is a centralized management capability introduced by Broadcom in ware Cloud Foundation (VCF) in version 9.0 to deploy, scale, patch, and operate multiple VCF instances. It encompasses multiple components, including VCF Operations, VCF Automation, vCenter, NSX Manager, vSphere Cluster, and workload domains.
QUIRSO said there is no overlap between this activity and the chain of events involving the abuse of CVE-2026-59310 on the same system starting August 3, adding the newly created "vcenter_admin" administrator account was not used in subsequent phases of the attack.
Exploitation of CVE-2026-59310
As for the exploitation of CVE-2026-59310, the first activity involved the cron daemon (aka crond) logging a malformed cron file called "zz-poc59310-syslog.log." In the next step, a curl command (or alternatively a wget command) is executed to retrieve a backdoor from "5.34.177[.]38:9861" and execute it, and then remove the log file.
The naming convention of the log file is significant as it is a direct reference to the CVE identifier and that it was a proof-of-concept (PoC) devised after details of the flaw became public knowledge.
"The '-syslog.log' suffix also mirrors the vCSA remote syslog file naming convention, but the file appears under /etc/cron.d rather than the configured syslog output directory," QUIRSO explained. "This suggests that the vCSA syslog server was abused to place files in a privileged execution location. While some files were malformed and not executed by cron, at least one file successfully executed and placed the 'linuxFile' backdoor on the system."
The linuxFile implant is designed to provide remote command execution capabilities to the attacker. It establishes a connection to its controller over a WebSocket channel to receive instructions, executes them through /bin/sh, and transmits the results back to the attacker.
"The C2 [command-and-control] address is XOR-obfuscated and decoded at run-time, while communications are protected using the malware's own application-layer cryptography despite using an unencrypted ws:// transport," Szadkowski told The Hacker News via email. "It also automatically reconnects on failure and contains routines for establishing persistence through systemd and cron."
The threat actor behind the operation also relied extensively on cron to execute malicious payloads, including to fetch and run a shell script ("esxi.sh") from the IP address "185.144.28[.]120:3232." The shell script then serves as a downloader and persistence installer for an architecture-specific reverse SSH ("reverse_ssh") binary that's retrieved from the same infrastructure.
Other cron jobs related to creating staging directories, downloading executables, changing their permissions, and running them, while referencing servers at "192.255.141[.]13:8080" and "5.34.176[.]100:5244." In what appears to be an operational security blunder, the latter has been found to expose the reverse SSH binaries toolset via an AList directory listing.
A brief description of some of the various actions carried out by the threat actor is as follows -
- Deploying "linuxFile" (aka systemlog or linux_x86), which connects to "ws://intel.se9ly9upbhay.shop:8080/ws" and establishes persistence via a systemd service.
- Setting three cronjobs impersonating legitimate VMware services: vmware-vpxd-stats-* (facilitates an SSH-based remote access channel by adding the attacker's SSH public key to the authorized keys file), vmware-perf-collect-* (drops a JSP web shell named "vmware-perf-update.jsp"), and vmware-perf-sync-* (drops the same web shell and runs a Base64-encoded script that performs credential access and sets up a new account called "adminuser," which is then added to the vSphere SSO Administrators group.
- Creating two additional accounts: adding "vcadmin" to vSphere with a Base64-encoded Python script dropped on disk via bash commands run in a cronjob and creating a vSphere admin account via an external LDAP "Add" operation against vCenter's VMware Directory Service (vmdir) from a remote client by using a pre-existing but compromised administrative account.
- Creating a file named "/etc/sudoers.d/vmware-perf" with a configuration that grants the "perfcharts" service account unrestricted, non-interactive passwordless sudo access to root.
- Running shell scripts like "/tmp/.vmware-perf-upd.sh" to obtain credentials for vmdir by querying the HKEY_THIS_MACHINE\services\vmdir registry location. If this method fails, it searches for VMware's vmafd Python module and calls GetMachineName(), GetMachinePassword(), and GetDomainName() to get the distinguished name and password associated with the vCenter machine account. The stolen credentials are used to conduct privileged directory modifications, including adding the aforementioned "adminuser" identity to the Administrators group.
- Using vSphere API to perform discovery operations and "esxi.sh" to deploy the reverse_ssh client.
- Creating local accounts on the ESXi hosts (e.g., "adminuser") to enable ransomware encryption.
- Taking steps to evade detection, reduce forensic visibility, and blend into the VMware environment.
The attack ultimately paves the way for the deployment of a ransomware on ESXi hosts that encrypts files with the ".babyk" extension, which is typically associated with Babuk-derived ransomware. It's not clear if this was the end goal of the campaign, or if the Babuk-derived payload was "selected opportunistically or even intentionally" to confuse attribution efforts.
QUIRSO told the publication it cannot assess at this stage if the ransomware strain was deployed across other compromised systems as the analysis was limited to only one of the infected systems. However, based on the investigation so far, it's suspected that the deployment of the locker may not have been the primary objective of the campaign.
Szadkowski likened the deployment to a smokescreen engineered to distract defenders from the main intrusion and thwart analysis by encrypting the ESXi log files, thereby preventing access to telemetry data that could have offered more insights into threat actor activity.
"Exploitation of CVE-2026-59310 provided the actor with immediate, non-interactive code execution in a root context on the vCenter Server appliance," the researchers said. "Subsequent commands recorded by CROND were therefore already being executed as root, giving the actor unrestricted access to the underlying VCSA without first having to compromise an unprivileged local account and escalate from it."
from The Hacker News https://ift.tt/dM4zODG
via IFTTT
North Korean Remote Workers Are Infiltrating Government and Businesses: How to Expose Them Before Hiring
Companies are used to thinking about attackers as outsiders trying to break in.
North Korean IT workers flip that model. They apply for jobs, pass interviews, receive legitimate credentials, and can end up inside the same systems companies spend millions trying to protect.
That risk is no longer theoretical. The FBI is now investigating a North Korean remote IT worker who reportedly worked for a U.S. federal agency.
For CISOs, the priority is clear: spot the warning signs before a fraudulent hire becomes trusted access.
When the Threat Gets Hired
A recent joint investigation by Mauro Eldritch (BCA LTD), Heiner GarcÃa (NorthScan), and ANY.RUN showed what this looks like from inside the operation.
Researchers deliberately hired suspected DPRK developers linked to Lazarus Group and gave them what looked like ordinary virtual desktops. In reality, they were controlled ANY.RUN Sandboxes, capturing their activity in real time.
The operation exposed forged identities, remote-access tools, AI-assisted workflows, and VPN and VPS infrastructure.
See the full investigation on the ANY.RUN blog for the recorded interviews, live operator activity, infrastructure findings, and complete toolset breakdown.
The Red Flags Start Before Day One
The investigation showed that the strongest warning signs were often small inconsistencies across the hiring process rather than one obvious giveaway.
Security and hiring teams should pay closer attention to:
- Identity details that don’t line up: addresses, states, documents, or banking information that contradict each other.
- Signs of document manipulation: unusual metadata, visual inconsistencies, or evidence that an ID has been altered with AI.
- Interview behavior that feels assisted: repeated off-screen glances, delayed responses, or dependence on live translation and AI tools.
- Location mismatches: network activity that does not match where the candidate claims to live or work.
None of these signals proves malicious intent on its own. But when several appear together, they should trigger deeper verification before the candidate receives company access.
How CISOs Can Keep a DPRK Operative Off the Payroll
The Famous Chollima investigation showed that there is rarely one obvious sign that gives a fraudulent worker away. Instead, the clues appear across identity documents, interviews, location data, infrastructure, and activity after onboarding.
Based on what researchers observed, here are several steps security leaders can take to make it much harder for a spy to get onto the company payroll.
1. Fully Verify the Person Behind the Documents
A convincing identity document should not be the end of verification.
The researchers encountered manipulated IDs, stolen identities, conflicting personal information, and financial details that did not always match the person being hired.
For sensitive remote roles, CISOs should make sure identity checks use several independent signals. The candidate’s documents, location, interview behavior, employment history, and financial details should tell a consistent story before access is approved.
Roles with access to source code, cloud infrastructure, production systems, or financial assets should receive a higher level of scrutiny from the start.
2. Give Security Teams a Safe Way to Validate Suspicious Activity
The researchers used specially configured ANY.RUN Sandbox environments to observe the operatives' activity without exposing real corporate systems. This gave them visibility into the files they opened, tools they used, network connections they made, and other behavior that would have been difficult to assess from identity checks alone.
CISOs can apply the same principle by making sure security teams have access to interactive sandboxes like ANY.RUN when suspicious files, links, scripts, or tools appear around employee activity.
Instead of relying only on alerts or isolated indicators, teams can safely examine how the activity behaves and gather stronger evidence before deciding whether escalation or containment is necessary.
This can help reduce uncertainty, speed up response, and lower the risk of suspicious activity reaching critical systems.
3. Check Whether the Same Infrastructure Appears in Your Environment
The investigation uncovered specific infrastructure used by the suspected DPRK operatives. Security teams can cross-check these indicators against historical logs, EDR telemetry, proxy records, DNS data, and other security sources to see whether the same infrastructure has already appeared inside the organization.
Examples from the investigation include:
- IPv4: 62[.]33[.]223[.]165 // INVESTSTROY-NET (InvestStroyTrest)
- IPv4: 89[.]187[.]185[.]11 // DPRK-operated VPS
- IPv4: 45[.]77[.]71[.]42 // DPRK-operated VPS
- IPv4: 185[.]152[.]67[.]39 // DPRK-operated VPS
- IPv4: 104[.]250[.]148[.]58 // AstrillVPN exit node
- IPv4: 192[.]200[.]115[.]226 // AstrillVPN exit node
- IPv4: 107[.]150[.]38[.]250 // AstrillVPN exit node
- IPv4: 206[.]217[.]134[.]34 // AstrillVPN exit node
- IPv4:199[.]168[.]112[.]175 // AstrillVPN exit node
- 0x8953B9661339a48f4E6408aA1B359CD49F3A6CAd
- 0xA3D6938f152C47A411263573Bb3AF324C25A8eba
- 0xB26A7C7EA6D75956EbD8c5D294524903b1cf13D0
A match should not be treated as proof of DPRK activity on its own, but it can be a strong reason to look deeper when combined with other suspicious signals.
With ANY.RUN’s Threat Intelligence Lookup, security teams can investigate indicators in more context, including related sandbox sessions, associated threats, and targeting patterns across countries and industries.
This helps teams determine whether an unusual connection is isolated or part of a broader malicious pattern and prioritize the cases that need attention first.
4. Turn Investigation Findings into Ongoing Detection
The researchers identified infrastructure used by the suspected DPRK operatives, including IP addresses, VPN endpoints, and VPS providers.
For CISOs, the next step is making sure findings like these are not checked once and forgotten. They should become part of ongoing detection, so security teams can spot the same or related infrastructure if it appears elsewhere in the environment.
ANY.RUN’s Threat Intelligence Feeds can support this by continuously supplying fresh indicators from real-world investigations to existing security tools.
That turns intelligence from cases like this into earlier warning signs for future activity.
Don’t Let a Fraudulent Hire Become a Trusted Insider
North Korean remote-worker schemes show why hiring can no longer sit outside the security conversation. A candidate may pass interviews, present convincing documents, and receive legitimate access long before traditional security controls see anything suspicious.
For CISOs, the priority is to reduce that gap: verify identity more deeply, give security teams the right solutions to validate suspicious activity, check known infrastructure against the environment, and turn confirmed findings into ongoing detection.
Give your security team the visibility and context to investigate suspicious activity faster and contain threats before business impact grows.
Found this article interesting? This article is a contributed piece from one of our valued partners. Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.
from The Hacker News https://ift.tt/TDaMvx8
via IFTTT
Sunday, August 16, 2026
Will OSS Models Take Over?
SUMMARY: This episode is the second part and explores the flip side of OSS models. Last episode, we discussed the potential decline; this episode, we’ll talk about the potential positive future of OSS models. Aaron and Brandon explore the future of open source AI models, the role of industry consortia, and how major tech companies like NVIDIA, Apple, and Google are shaping the AI landscape. They discuss the potential for open models to become industry standards and the strategic motivations behind these moves.
SHOW: 1054
SHOW TRANSCRIPT: The Enterprise AI Show #1054 Transcript
SHOW VIDEO: https://youtu.be/w238Y1ZKG1Q
SHOW SPONSORS:
Topic: Are we seeing the end of OSS models?
- Why now? NVIDIA Open Secure AI Alliance (all except Anthropic joined) & Linux Foundation is managing proposals
- Past: OSS runs the world… Up until now, there hasn’t been an overarching “AI Model” project managed by the CNCF or Linux Foundation that has gained any traction
- Present:
- As model sizes increase, who pays for training? I think the DB market is the closest parallel here, and it's also where the most OSS rug pulls have happened in the past. Is this history repeating itself, but also a lesson learned because so many DB companies got burned?
- Future: Someone will have to donate a trillion+ parameter model to a foundation. My bet is NVIDIA will eventually drive this through Nemotron; it makes the most sense, and they have the most to lose if OpenAI and Anthropic take over and also eventually use their own chips.
FEEDBACK?
- Email: show @ the enterprise ai show dot come
- Bluesky: @TheEntAIShow.bsky.social
- Twitter/X: @TheEntAIShow
- Instagram: @TheEntAIShow
from The Cloudcast (.NET) https://ift.tt/bFwRve9
via IFTTT
Friday, August 14, 2026
Healthcare IT infrastructure: how to reduce downtime
IT infrastructure failures in healthcare aren’t measured in downtime percentages. They are measured in canceled appointments, delayed care, and halted production lines. Standard high-availability approaches can break down here because the environment is uniquely unforgiving, with complex interface chains such as Health Level Seven International (HL7) and Fast Healthcare Interoperability Resources (FHIR), strict regulatory oversight, and no IT staff on site at many remote locations.
When a recovery plan misses the mark, a simple hardware hiccup can quickly become a serious business problem. Clinics face missed billing windows and idle staff on payroll, while patients experience frustrating delays that can damage trust and push them toward competitors. Let’s break down how to engineer true resilience for clinics, diagnostic labs, imaging centers, and manufacturing floors where server downtime is simply not an option.
What is healthcare IT, and why does uptime matter?
Healthcare IT is the technology chain connecting people and equipment to data. When a clinician opens an Electronic Health Record (EHR), the request depends on the entire transaction path, from identity and DNS to storage and network. A lab analyzer similarly relies on middleware and interface engines before results reach the Laboratory Information System (LIS). If any link fails, the service can remain technically online while still being unusable.
Cloud hosting changes where components run, but the dependency chain remains. A SaaS EHR still requires local internet access and authentication, while a cloud archive needs a Digital Imaging and Communications in Medicine (DICOM) gateway and reliable bandwidth. Patient-facing touchpoints like mobile apps and portals extend the chain further, requiring authentication and duplicate handling before data reaches the patient record.
Applications typically exchange this data via HL7 standards, with HL7 v2 used for clinical interfaces and FHIR providing web-oriented APIs. Both provide structure for data exchange, but neither guarantees correct patient matching or a functioning interface.
For downtime planning, workloads generally fall into three buckets: those that stop operations immediately, those that allow a degraded manual workflow, and those that can queue data for later synchronization. This behavior determines the recovery target and helps you decide which systems need local failover and which can wait for a standard restore.
Core components of healthcare IT infrastructure
A resilient design starts with understanding what each infrastructure layer needs during normal operation and during a failure. As you build or review the environment, look at each layer from the perspective of a failure and ask what the rest of the workflow depends on.
Network paths need to be defined by failure domain, not convenience. When you map these paths, separate clinical access, management traffic, storage replication, and backups so they do not compete for the same pipe. Segment networks to limit lateral movement, encrypt data in transit, and log everything that touches protected health information (PHI).
Compute includes physical hosts, hypervisors, and VMs running the EHR, interface engines, communications, and specialist systems. When you size HA, assume that one host is unavailable and verify that the surviving host can carry the priority workload on its own.
Storage includes primary volumes, synchronized replicas, archives, and backups, each serving a different recovery purpose. When you review your storage design, make sure replication is complemented by versioned or immutable copies. Encrypt data at rest, because live replication alone does not provide a reliable recovery point if corrupted or encrypted data is replicated to the other copy.
Identity and access management determines who can see or change clinical data. This includes directory services, role-based permissions, service accounts, multifactor authentication for administrators, and controlled emergency access.
Monitoring and auditing cover the remaining operational requirements. Infrastructure telemetry tracks latency, capacity, failed paths, and resynchronization load. Audit logs record access and changes for incident review and regulatory evidence.

Figure 1. Core components of healthcare IT infrastructure.
On-premises, cloud, or hybrid?
The deployment model affects the failure domains you need to account for. On-premises infrastructure keeps device-facing services local and reduces WAN dependence. Cloud services reduce local hardware management, but provider availability, identity, connectivity, and data location become part of the failure model. Hybrid designs commonly keep equipment gateways on site while hosting EHR, analytics, or archives elsewhere.
When you compare these models, look at where each dependency sits and what happens when connectivity to that location is lost. Deployment location does not establish HIPAA compliance. Under the U.S. Department of Health and Human Services (HHS) cloud guidance, a provider handling ePHI on behalf of a covered entity or business associate is a business associate, and a compliant Business Associate Agreement (BAA) is required. Selection should therefore follow recovery targets, WAN tolerance, data-location rules, and operational responsibility.
Healthcare IT use cases by industry
The same uptime target can lead to very different infrastructure designs depending on the workload. To see how these differences affect infrastructure decisions, you can map each healthcare environment to the workflow at risk, the dependency teams often overlook, and the design decision that follows.
| Healthcare environment | Critical workflow | Downtime impact | Design focus |
|---|---|---|---|
| Clinics and hospitals | EHR access, medication, and orders | Staff lose access to clinical workflows | Protect identity and interfaces; size for peak load |
| Diagnostic laboratories | Specimen-to-result path | Results cannot be matched or reported | Test analyzer buffering and the full interface path |
| Medical imaging centers | PACS ingest, retrieval, and reporting | Studies queue or prior images disappear | Size for ingest, archive growth, and resynchronization |
| Medical research | Study data and audit records | Data entry stops or record state is uncertain | Use controlled recovery with retained audit trails |
| Pharmaceutical manufacturing | MES, batch records, and quality status | Production or batch review pauses | Validate failover and infrastructure changes |
| Distributed and remote sites | Local clinical services | Staff wait for remote repair | Standardize two-node deployments and monitoring |
The table highlights an important point: recovering the server is only one part of restoring the workflow. The actual recovery target should be tied to the business or clinical process that depends on that server.
A laboratory recovery test does not end when the LIS server boots back up. It ends when a test result successfully maps to the correct patient MRN. When you test a laboratory recovery scenario, measure the buffer depth against your actual test volume: how many hours of data the analyzer can hold locally before overwriting data or overwhelming the interface engine with a large backlog.
Imaging represents a heavy storage workload because a single node often handles three concurrent jobs: serving active studies, ingesting new ones, and rebuilding a replica after failover. This simultaneous load can saturate disk IOPS and expose limitations that daily-average capacity figures do not show. If you size the environment using only daily-average figures, you can easily underestimate the load during failover.
Research and pharmaceutical systems introduce another compliance requirement: recovery must preserve a controlled record state. Under FDA guidelines for electronic records and data integrity, failover evidence matters just as much as raw recovery speed. You need to be able to demonstrate what happened during recovery and verify that the resulting records remain trustworthy.
At distributed sites, architecture is closely tied to supportability. If you manage multiple remote locations, standardization becomes especially important because your central team needs to troubleshoot the same architecture remotely. A standardized cluster and monitoring setup allows a central team to manage dozens of locations, while a unique stack at every site makes remote diagnostics, maintenance, and spare-parts planning much harder.
What causes healthcare IT downtime?
The failed component is not always the obvious server. Identity and DNS outages can make several healthy applications inaccessible at once. A bad virtual-switch change can isolate every VM on a host. Storage latency can look like an application problem. Interface engines, certificate services, and database connections often sit outside the workload inventory used for HA planning.
Maintenance introduces its own failure modes. An update can change a driver, a firewall rule can block an interface, or a certificate can expire after the team has tested only the application login and not the interfaces behind it.
When you review your HA plan, make sure these dependencies are included in the failure model, not just the servers and VMs. After any change in the EHR path, check the complete workflow:
- Can a clinician open a chart and place an order, not just reach the login screen?
- Is the interface engine routing messages end to end, not just showing a green “connected” status?
- Did any firewall or network access control list (ACL) change block a port that an interface or service account depends on?
- Are certificates still valid on every hop, not only the one the team remembered to check?
- Does the failover path still complete within the RTO when tested with production-representative load?
Cyberattacks and site loss create different recovery problems. Ransomware can damage every synchronized copy that is reachable with the same credentials. Fire, flood, or a prolonged power failure can remove both cluster nodes at once. Those events move the response from local failover to backup or disaster recovery.
Why is backup not enough?
Backup cannot take over a live workload when a host fails. The data must be restored to working compute, the application must start, and its dependencies must reconnect. That recovery process can take longer than a clinical or production workflow can tolerate.
HA and backup address different failure scenarios, and treating them as interchangeable creates gaps in the recovery plan. When you review your recovery strategy, it helps to look at what each layer can actually recover from:
| HA (synchronous replication) | Immutable / air-gapped backup | |
|---|---|---|
| Protects against | Host, disk, or network hardware failure | Ransomware, accidental deletion, data corruption |
| Recovery point | Seconds (current state) | Depends on backup schedule (hours to a day, typically) |
| Recovery time | Automatic failover, usually minutes | Manual restore, potentially hours |
| What it can’t stop | A bad or malicious write replicates to both copies instantly | A slow-burn hardware failure over weeks |
Synchronous replication preserves the current state, including an accidental deletion, a corrupt write, or ransomware encryption. Two healthy replicas are therefore not two recovery points. Critical workloads still need backups isolated from production credentials and failure domains, on media that the HA cluster’s own admin credentials can’t access.
RTO should include the full service recovery, not just the time required to copy data. RPO should reflect the point to which the application can be returned without breaking its workflow or record obligations. Restore order matters too: identity and DNS may need to return before the database, followed by the application and its interfaces.
This is why a backup test can look successful while the actual service recovery fails. Testing one VM in isolation may produce a clean backup report, but it does not prove that the complete application stack can be restored and brought back into service within the required RTO.
How does a two-node HA cluster work?
In a two-node design, each server provides compute and local storage. Synchronous replication maintains matching data copies across the nodes and presents them as shared storage to the virtualization cluster. If one host fails, its VMs restart on the surviving host using the available storage copy.
The replication network becomes a critical part of this design. Synchronous writes add a network dependency because both copies need to commit the write before it is acknowledged. As a rule of thumb, sync replication between cluster nodes requires round-trip latency in the low single-digit milliseconds, commonly around 2-3 ms or better. As latency increases beyond that range, write latency on every VM can also increase, including during normal operation.
Replication links need to maintain that low latency consistently, along with enough bandwidth for production traffic and resynchronization after a node returns. When you size these links, account for the rebuilding state as well as normal workload traffic.
Quorum logic prevents both nodes from serving conflicting copies after communication breaks. The exact mechanism may use heartbeat channels, a witness, or node-majority logic. Where the witness lives matters more than many designs account for. If you place it in the same rack, on the same power circuit, or on the same physical host as one of the two nodes, a single event, such as a tripped breaker or rack switch failure, can take out a node and the witness together.
The surviving node can then lose quorum as well, and a cluster that should have failed over may stop with otherwise healthy hardware sitting idle. When you design a two-node cluster, treat witness placement as part of the failure-domain design.
Capacity is set by the failed state, not the normal one. One node must run the priority VMs and still meet application response targets. A technically successful failover that leaves Picture Archiving and Communication System (PACS) retrieval or LIS transactions too slow hasn’t met the operational requirement, regardless of what the uptime dashboard reports.
Failure mechanics: quorum, storage, and network design
Network partition and split brain
In a two-node cluster, a failed heartbeat or replication path can leave both servers powered on but unable to see each other. If both sides continue serving the same workload, their storage state can diverge. Quorum prevents this outcome by allowing only the partition with a majority of votes to keep clustered workloads online.
The witness holds a vote, not clinical data, which is why its placement is a network and power design decision. iSCSI, SMB Direct, and RDMA carry storage traffic; none of them determine quorum. Forcing both partitions online defeats the protection and can create conflicting writes.
When you plan the network, keep the quorum path and storage path in mind separately. A high-performance storage network does not compensate for poor quorum design, and a healthy witness does not fix an overloaded replication link.
PACS: archive size, burst traffic, and cache
DICOM is structured, but PACS storage behaves more like a large-object workload than a conventional transactional database. Daily acquisition may be predictable. A bulk migration, archive restore, or node resynchronization rarely is and can introduce a much larger sustained stream.
With synchronous replication, a foreground write is acknowledged only after both copies commit it. If a PACS import shares disks or links with EHR and LIS workloads, a large transfer can raise latency across the whole cluster. Write-back cache can absorb a short burst, but once the cache fills, throughput falls to the rate of the slower replica or network path.
For this reason, capacity tests should combine clinical reads, peak image ingest, and full resynchronization instead of benchmarking each workload in isolation. Keep recent or frequently read studies on the active HA tier. Move older DICOM objects to a PACS- or VNA-supported archive tier when retention and retrieval targets allow it.
When you plan archive migrations, throttle bulk transfers and verify that the archive remains searchable before removing the source copy. This gives you a practical recovery and availability check.
Storage and replication network
Separate client, management, cluster, storage, and backup traffic using physical paths, dedicated VLANs, and QoS. For iSCSI, use dedicated adapters or HBAs with MPIO. If iSCSI crosses a router, thoroughly test latency, packet loss, MTU, and failover across the entire path.
For a busy two-node storage cluster, 10GbE is a practical starting point, while 25GbE becomes appropriate once PACS ingest, VM writes, and resynchronization traffic saturate lower speeds. Always size links for the failed or rebuilding state rather than average daytime utilization.
Jumbo Frames require matching MTUs on every NIC and switch. SMB Direct needs RDMA-capable adapters and SMB Multichannel, while RoCE requires appropriate DCB and PFC configuration. Always run load tests with active replication and monitor latency, retransmissions, dropped packets, and queue depth.
The important part is to test these conditions together. A network can perform well under a synthetic bandwidth test and still struggle when VM writes, PACS ingest, and storage resynchronization compete for the same resources.
External SAN vs. 2-node HCI
For local HA storage under a healthcare workload, organizations typically choose between a dedicated SAN array and synchronous replication across local server disks, without a separate storage array.
A SAN can fit a large data center with dedicated storage admins, but it becomes more difficult to operate across a distributed clinic network. Each site with an external array introduces specialized hardware, firmware, support contracts, and additional failure modes into locations that may not have dedicated IT staff.
Two-node HCI collapses this complexity. Each server carries its own storage, replicates synchronously to its pair, and presents shared storage to the hypervisor without a separate SAN appliance or storage team.
StarWind implements this model in two ways. Virtual SAN (VSAN) runs on existing server storage for teams with in-house management skills. Alternatively, HCI Appliance (HCA) delivers pre-integrated compute, storage, virtualization, and support as a single box, ideal for multi-location rollouts where standardization outweighs hardware reuse.
Ultimately, the choice for distributed clinics comes down to operational overhead. Reusing servers can save money upfront, but it can also require more engineering time later when teams have to troubleshoot different hardware and software stacks across multiple sites. Standardized hardware with a single vendor escalation path can be easier to operate and support over a multi-year deployment.
When you compare the options, look beyond the initial hardware cost. Consider who will troubleshoot a failed component at 2 a.m., how quickly replacement parts can reach a remote site, and whether the central IT team can apply the same recovery procedure at every location.
Regardless of the storage approach, the protected unit remains the virtualized workload on the cluster. WAN services, SaaS availability, medical devices, and site-level disaster recovery remain separate design considerations.
How to control healthcare IT infrastructure costs
Cost control starts by matching protection to the workload instead of buying the same availability tier for every application. HA belongs where the operational cost of waiting for a restore exceeds the cost of redundant infrastructure. Lower-priority systems can use standard backup and recovery.
When you evaluate existing infrastructure, look at the failed-state requirements first. Existing servers are economical only if they can carry the failed-state load and have enough supported life left to justify the integration work. Cloud estimates also need to include redundant connectivity, data transfer, backup, monitoring, security tooling, and support, not just the compute line item.
Across remote sites, repeatable configurations and central monitoring can save more over time than a small discount on one-off hardware. If you manage multiple locations, standardization also reduces the time your team spends troubleshooting different configurations and planning replacement parts.
Healthcare security and regulatory requirements
Regulations do not prescribe a two-node cluster or a particular storage product. They shape the controls and evidence around the systems that store or process regulated records. Your infrastructure design therefore needs to support the required security and compliance controls without treating the HA platform itself as proof of compliance.
For covered entities and business associates, the current HIPAA Security Rule requires safeguards for the confidentiality, integrity, and availability of electronic protected health information. Its contingency-planning provisions cover backup, disaster recovery, emergency-mode operations, and plan testing. HA can support the availability component, but the compliance record comes from the broader risk and control program.
The CMS Emergency Preparedness Rule applies only to designated Medicare- and Medicaid-participating provider and supplier types, with requirements that vary by category. The CLIA program focuses on accurate, reliable, and timely human laboratory testing. For FDA-regulated research and manufacturing, Part 11, CGMP, validation, and data-integrity requirements may apply, depending on the system and records involved.
In you operate in the EU, GDPR governs personal-data processing, and NIS2 may impose cybersecurity and incident-reporting duties on healthcare entities within scope. Software that qualifies as a medical device falls under the EU MDR; IEC 62304 provides a lifecycle-process standard for medical device software. These requirements do not apply to every healthcare application or provider, so you need to check the applicable scope and obligations case by case.
Encryption keys and backup isolation
Security planning should continue down to the key-management level. An architecture record should clearly separate encryption at rest from encryption in transit. For a Windows cluster using BitLocker-protected volumes, key protectors sit at the volume and cluster layer, and recovery material must be escrowed outside the two-node failure domain. If replication uses TLS or IPsec, session keys exist for the connection, while long-lived credentials remain in the protected host keystore or an external KMS or HSM selected by the platform.
Your design should document where these keys and credentials are stored, who administers them, how rotation and revocation are handled, and what happens during failover if the key service becomes unavailable. NIST SP 800-57 provides the key-management framework; the exact implementation remains product- and hypervisor-specific.
Backup isolation is equally important. A two-node cluster does not provide an air gap or an immutable backup. Both replicas are online and receive the same authorized writes, so ransomware, deletion, or corruption can reach both. Cyber-recovery copies need a separate security and failure domain, separate credentials, and immutable retention or offline media.
CISA recommends offline backups, and NIST distinguishes replication from immutable and point-in-time recovery. When you test recovery, verify that the isolated copy can restore the complete application workflow, not just a single file.
An infrastructure review should connect each technical control to evidence: approved architecture, access rules, change records, backup results, failover tests, restore tests, and assigned owners. Product redundancy supports this work, but the surrounding processes and evidence still need to be maintained.
Conclusion
Planning for healthcare downtime means looking at workflows instead of isolated servers, whether the workload involves clinic access, lab results, or imaging queues. A two-node cluster can be a good fit when a failed host can’t wait for a rebuild, a single node can carry the operational load, and site recovery is managed separately.
As you plan your environment, define the protection boundary clearly. Use local HA to protect against on-site hardware and storage failures, while backups, network resilience, and documented recovery procedures address the failure scenarios that HA cannot cover. You need to keep the workflow running and make sure the recovery process works when the infrastructure is under pressure.
Frequently asked questions
How often should failover be tested in a clinical environment?
Test after any material infrastructure or application change and on a regular, risk-based schedule. A test is only complete when clinicians or staff can successfully finish the affected workflow, not simply when the virtual machine boots back up.
Who should participate in a healthcare recovery test?
At a minimum, IT, the application owner, and clinical, laboratory, or manufacturing staff who rely on the system daily should participate. Include compliance or security teams whenever regulated records or cyber-recovery scenarios are part of the test.
What belongs in a healthcare downtime runbook?
Include clear steps for who declares the outage, how clinical and operational workflows continue during the disruption, the exact order in which critical services must return, and how queued data or paper records will be reconciled once systems are back online.
Your runbook should also identify the people responsible for each step and the escalation path if recovery takes longer than the defined RTO. Keep it specific enough that someone who does not normally operate the environment can follow it during an incident.
from StarWind Blog https://ift.tt/hyHsYJO
via IFTTT