Wednesday, August 5, 2026

The hidden economics of healthcare IT: where cost accumulates

Healthcare organizations don’t experience cost as a single line item. It materializes in delayed workflows, constrained clinician time, and increasing operational load across already stretched teams.

Gartner notes that in 2026, healthcare provider CIOs face significant pressure to deliver digital value despite constrained IT budgets and recommends investing in initiatives that improve IT performance to generate savings rather than simply cutting spend (Gartner, 2026).

The largest drivers of cost in healthcare IT environments are not introduced at procurement. They emerge across three pressures CIOs face:

  • The rising capital cost of endpoints
  • The scarcity of clinical and IT labor
  • The operational cost of defending an increasingly distributed security surface

The sections that follow examine how each of these pressures materializes in day-to-day operations and how platform design either compounds or contains them.

The rising capital cost of endpoints

Endpoints in healthcare are inseparable from clinical workflows. Devices must continuously support EHR access, imaging and diagnostic integrations, and real-time data entry at the point of care.

At hospital scale, a 17% year-over-year increase in PC pricing is not an inconvenience, it is a capital reallocation (Gartner, 2026). That cost competes with capital allocated to imaging equipment, bed capacity, and clinical hires. When endpoint refresh consumes a larger share of capital budgets, it crowds out the clinical investments those budgets are also expected to fund.

A Forrester Total Economic Impact study, commissioned by Citrix, quantified the operational and financial outcomes healthcare organizations realized after deploying Citrix DaaS. The study illustrates what decoupling performance from the endpoint looks like in practice.

Citrix was deployed across compute-intensive clinical functions including patient care, electronic medical records, and clinical decision support, with compute and application workloads executing centrally rather than on the device. Instead of delivering performance, the endpoint provides access to the centralized environment where compute and applications actually run. That is why the devices supporting that workload did not need to be uniform or current.

One Forrester-interviewed health system extended this further by integrating badge login with its Citrix workstations, removing authentication friction from the same devices that no longer constrained performance (Forrester, 2026).

Refresh decisions can then be paced rather than forced, and capital can remain directed toward investments that scale clinical output.

The scarcity of clinical and IT labor

Workforce shortage is no longer a future risk in healthcare. It is a current operating condition. The AAMC projects a shortage of up to 86,000 physicians by 2036 (AAMC, 2025) and Becker approximates a shortage of 100,000 critical healthcare workers within two years (Becker’s Hospital Review, 2025). In that environment, every minute clinicians lose to system friction is more expensive than it was a year ago.

That friction is measurable. Becker’s Hospital Review found that in 80% of healthcare organizations, fewer than 70% of clinicians reported that their EHR responded quickly and 35% of nurses said they spend three or more hours per week on duplicative or unproductive documentation (Becker’s Hospital Review, 2025).

The Forrester study quantifies the impact of that friction on clinical productivity. Before improving application delivery, one healthcare organization reported latency and data synchronization issues across clinical systems. After stabilizing performance and aligning resources with demand, productivity increased by 30% (Forrester, 2026).

The same pattern appears in session access, where the friction is not only login time but the loss of clinical continuity every time a session ends. Clinicians restart their workflow at every new workstation; reauthenticating, reloading applications, and reopening patient records, with each transition taking 25 to 30 seconds.

Roaming virtual desktops remove that break in continuity. The clinician’s session persists across workstations, so they reconnect to the same live environment in five seconds or less, with their applications, records, and context already open (Forrester, 2026). Recovered across thousands of interactions per shift, those seconds compound into meaningful clinical capacity.

IT teams face the same structural gap. Specialized expertise is increasingly difficult to retain, and health systems cannot hire their way out of operational complexity. Citrix reduces that complexity by centralizing management, standardizing the delivery environment, and eliminating the configuration drift that generates most recurring tickets in distributed fleets.

When every user connects to the same managed environment, most of the conditions that trigger Level 2 and Level 3 escalations disappear at the source. After implementing Citrix DaaS, one healthcare organization eliminated more than 90% to 95% of those issues, cutting ticket volumes dramatically (Forrester, 2026).

The leverage that follows is significant. A composite healthcare organization supports 15,000 users on a team of ten Citrix engineers, with deployment spanning clinical staff, hospital staff, and corporate knowledge workers (Forrester, 2026).

The implication is not that fewer people are required. It is that the people who are there can focus on work that compounds.

The operational cost of defending a distributed security surface

Healthcare organizations operate under strict regulatory requirements, including HIPAA. Security architecture must support these obligations across distributed care environments. In endpoint-heavy environments, controls must be replicated and maintained across many systems. This increases operational effort and introduces variability in enforcement.

The cost of that fragmentation is increasingly visible. Over 80% of stolen protected health information records in recent years have originated outside hospital systems, in third-party vendors, business associates, and non-hospital providers (American Hospital Association, 2025). At an industry level, healthcare data breaches now average $7.42 million per incident and take 279 days to identify and contain, making healthcare the costliest sector for breaches for the fourteenth consecutive year (IBM, 2025).

The economic problem with perimeter-based security is not only the cost of breaches when they occur, but also the cost of defending the perimeter when they do not. Every endpoint, every third-party connection, and every control point require independent maintenance, patching, and audit.

The Forrester study captures this dynamic from the IT team’s perspective. One healthcare organization consolidated application hosting into a single centralized data center supported by Citrix and eliminated several regional data centers in the process. Each eliminated regional data center removed a full set of network boundaries, identity endpoints, patching cycles, and audit obligations that previously had to be defended independently (Forrester, 2026).

Consolidating hosting compresses the security surface itself, not just the infrastructure footprint.

Conclusion

Healthcare IT cost in 2026 is not determined by what is purchased. It is determined by how systems respond to the three pressures shaping the sector:

  • The rising capital cost of endpoints
  • The sustained scarcity of clinical and IT labor
  • The compounding operational cost of distributed security

Each of those pressures sits outside the procurement conversation. Each of them is increasing, and each of them is influenced more by how technology is designed to operate than by what appears on its invoice.

Where time, accuracy, and continuity define care delivery, the cost of healthcare technology is a function of how it performs against the conditions the sector is operating under, not the conditions it was designed for a decade ago.

For more information about Citrix for healthcare, click here.


Sources: 



from Citrix Blogs https://ift.tt/ArWwzEc
via IFTTT

Governance Is a Developer Experience Problem

This is the third post of a 3-part series by Docker Captain Karan Verma. Catch up on Part 1: Your Laptop Is the New Production Environment and Part 2: Runtime Enforcement, Not Runtime Advice.

The conversation around AI governance often starts with security. That’s understandable. When autonomous systems can execute commands, access tools, and interact with production-adjacent environments, organizations naturally focus on risk. But after spending time thinking about agent workflows, I’ve become convinced that governance is about more than security. It’s also a developer experience problem.

The Trust Bottleneck

Most organizations don’t struggle to adopt new tools because the tools are incapable. They struggle because the organization doesn’t trust them yet. The history of software development is full of examples. Cloud adoption accelerated when organizations became comfortable with cloud governance. Containers accelerated when teams gained confidence in isolation and operational controls. CI/CD accelerated when organizations trusted automated deployment pipelines. The pattern repeats. Capability arrives first. Trust arrives later. Adoption follows trust. AI agents are no different.

image1 2

Caption: Capability alone does not drive adoption. Trust enables organizations to delegate work, expand usage, and realize productivity gains.

The Wrong Tradeoff

Governance is often framed as a choice between speed and control. Move fast and accept risk. Or add controls and slow everyone down. In practice, the most successful developer platforms rarely make this tradeoff. Instead, they create environments where developers can move quickly because boundaries already exist. A developer deploying through a mature platform doesn’t need to think about every networking rule, access policy, or infrastructure safeguard every time they ship code. The platform already provides those guarantees. The same principle applies to agent systems. The goal isn’t to force developers to manually approve every action. The goal is to create environments where useful actions can happen safely by default.

A Tale of Two Teams

Imagine two engineering teams using the same coding agent. The first team allows agent usage only in limited experiments because nobody is completely certain what the agent can access, execute, or modify. Every new workflow requires additional review. Every new capability triggers a discussion about risk.

The second team operates within clearly defined boundaries around execution, tools, and credentials. Developers understand where agents run, what systems they can access, and how activity is observed.

The underlying model is identical. The difference is trust. Over time, that difference may matter more than the model itself. Organizations rarely scale technology they do not trust.

Why Boundaries Create Freedom

This idea sounds counterintuitive at first. Boundaries feel restrictive. But in software systems, boundaries often enable autonomy rather than limiting it.

When organizations know:

  • where agents run,
  • what agents can access,
  • which tools agents can use,
  • how activity is observed,

They become more comfortable delegating work. Without those boundaries, every workflow becomes an exception process. Every deployment requires discussion. Every new capability triggers concern. Every new tool requires negotiation. Governance reduces uncertainty. Reducing uncertainty increases trust. And trust enables adoption.

The Platform Shift

One thing that stands out in recent discussions around agent infrastructure is that governance is increasingly moving into the platform itself. Developers shouldn’t need to become security experts every time they use an agent. Just as developers rely on platforms to handle identity, networking, deployment, and observability concerns, governance increasingly becomes part of the environment where agents operate. When governance is embedded into the platform, developers spend less time worrying about boundaries and more time focusing on outcomes. That’s a developer experience improvement as much as a security improvement.

Governance as an Enabler

The organizations that adopt agents most successfully may not be the organizations with the fewest controls. They may be the organizations with the clearest controls. Clear boundaries create confidence. Confidence enables delegation. Delegation unlocks productivity. Viewed through that lens, governance is not the thing slowing agent adoption. It is one of the things that makes large-scale adoption possible.

Looking Ahead

The conversation around AI agents often focuses on what models can do. Increasingly, I think the more interesting question is what organizations are willing to trust them to do. That trust won’t come from capability alone. It will come from visibility, accountability, and well-defined boundaries because the future of agentic software is unlikely to be determined solely by the most capable agents. It will also be shaped by the environments that make those agents trustworthy enough to use at scale.

Learn more



from Docker https://ift.tt/3mLdAaR
via IFTTT

Kali365 Weaponizes Microsoft Authentication Against US Companies: New Enterprise Risk

Kali365 is turning a legitimate Microsoft login into a gateway to corporate data.

The phishing kit targets US organizations with attacker-controlled device codes that victims approve on Microsoft's real authentication page. Once access and refresh tokens are issued, attackers may retain access to email, documents, and cloud resources, creating a direct path to data exposure, financial fraud, operational disruption, and costly incident response.

How Kali365 Targets US Organizations

Kali365 is a device code phishing kit built to abuse legitimate Microsoft authentication. ANY.RUN telemetry records more than 80 public sessions linked to the campaign each week, with the United States emerging as its main geographic target.

One of these sandbox sessions shows a SharePoint-themed lure used to draw the victim into the authentication flow.

View the analysis session and gather IOCs

SharePoint-themed Kali365 lure analyzed inside ANY.RUN’s Interactive Sandbox

Based on the research, the attack unfolds in three main stages:

Lure: The victim is presented with a page impersonating a trusted business service such as SharePoint, OneDrive, or DocuSign.

Microsoft authentication: The page redirects the victim to Microsoft's legitimate device login portal and asks them to enter an attacker-provided code.

OAuth access: Once the victim completes authentication, attackers may obtain access and refresh tokens that provide continued access to Microsoft 365 email, documents, and cloud resources.

Reveal the full phishing chain in as little as 60 seconds to reduce response delays and prevent a single compromised account from becoming a wider business incident.

Reduce Incident Risk

What Kali365 Can Cost the Business

A single approved device-code request can expand into a wider Microsoft 365 compromise. For US companies, the consequences may include:

  • Financial fraud: Compromised email accounts can support invoice manipulation, payment fraud, and business email compromise.
  • Sensitive data exposure: Attackers may access corporate email, internal files, customer information, and confidential documents.
  • Operational disruption: Unauthorized access to cloud services can interfere with daily communications and business processes.
  • Higher response costs: Fewer obvious phishing indicators can delay detection and make containment more complex.
  • Compliance and reputational risk: Exposure of regulated or customer data can trigger reporting obligations and damage trust.

As the victim authenticates on Microsoft's legitimate page, the activity may appear routine at first, giving attackers more time to misuse trusted access before the incident is confirmed.

Three Priorities for Reducing Kali365 Risk

Kali365 cannot be addressed through email filtering alone. Security leaders need current campaign intelligence, faster validation of suspicious activity, and better preparation for how the threat may evolve.

1. Expand Detection with Actionable Phishing Intelligence

Kali365 operators can rotate domains, URLs, and hosting infrastructure as campaigns evolve. Indicators from one confirmed case may quickly become outdated, leaving gaps across the rest of the environment.

Fresh phishing IOCs should reach SIEM, SOAR, TIP, firewalls, and other security controls where they can support alert enrichment, retrospective searches, and blocking decisions. ANY.RUN's Threat Intelligence Feeds deliver newly observed indicators through STIX/TAXII, API, and SDK.

Get fresh and trustworthy IOCs on emerging threats for deeper investigations

The intelligence is drawn from sandbox investigations submitted by more than 15,000 organizations and 600,000 security professionals worldwide. Each IOC links back to the session where it appeared, giving defenders the full context needed to verify the threat and identify related Kali365 infrastructure.

2. Give Tier 1 the Evidence Needed to Act on Kali365

As victims authenticate on Microsoft's legitimate device login page, Kali365 may look like normal activity at first. The real warning signs often appear earlier, in the lure, redirects, browser behavior, scripts, and attacker-controlled infrastructure.

ANY.RUN's Interactive Sandbox combines hands-on interaction with automated analysis to reveal the full attack chain faster, from the phishing page and redirect paths to network activity and the transition into Microsoft's authentication flow.

Tier 1 reports include AI summaries, recommendations and all the evidence needed for faster handoff

Auto-generated reports bring together the verdict, IOCs, TTPs, and behavioral evidence in a shareable format. This helps Tier 1 confirm malicious activity sooner, hand off complex cases with clearer context, and support faster containment before access spreads across Microsoft 365.

3. Turn Threat Research into Proactive Defense

Kali365 activity can be explored beyond a single alert by checking current campaign data in ANY.RUN's Threat Intelligence Lookup. The results provide context on related infrastructure, relevant sandbox sessions, lure screenshots, and targeting patterns.

For US-focused activity, teams can run the following query:

threatName:"kali365" AND submissionCountry:"US"

Kali365 activity targeting US organizations uncovered in ANY.RUN’s Threat Intelligence Lookup

The results show Kali365 activity across manufacturing, technology, healthcare, government, consulting, and MSSPs. This gives defenders a clearer view of where the campaign is active and which domains, URLs, and infrastructure may be connected to it.

Threat Intelligence Reports add a broader layer of preparation. These reports are manually compiled by ANY.RUN analysts and focus on active malware and phishing campaigns, including APTs and cybercriminal groups.

TI reports created by ANY.RUN analysts for deeper investigations

Each report includes investigation findings and TI Lookup queries that teams can apply to threat hunting, detection reviews, and incident enrichment. This helps SOC teams track emerging attack patterns earlier and prepare before similar activity reaches their environment.

Shut Down Token Abuse Before It Reaches the Business

Kali365 puts pressure on a part of the security stack many organizations still treat as trusted by default: cloud authentication.

The CISO challenge is to ensure the SOC can recognize when a legitimate login flow has been manipulated, trace the activity back to its source, and contain access before email, files, or business systems are affected.

Organizations using ANY.RUN have reported:

  • 94% faster threat triage, helping critical incidents move to action before they are delayed by alert backlogs.
  • Up to 21 minutes less MTTR per case, reducing the window in which attackers can expand access or misuse trusted accounts.
  • Up to 20% lower Tier 1 workload, creating more investigation capacity without immediately adding headcount.
  • 30% fewer Tier 1-to-Tier 2 escalations, allowing senior analysts to focus on complex incidents and higher-risk decisions.

These gains lower response costs, improve the use of existing SOC resources, and shorten the window for token abuse to escalate into fraud, data exposure, or operational disruption.

Contain identity-based threats with behavioral evidence before they reach critical business systems.

Cut MTTR by 21 Mins Per Case

Found this article interesting? This article is a contributed piece from one of our valued partners. Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.



from The Hacker News https://ift.tt/FUuJfvz
via IFTTT

Leaked n8n API Tokens Exposed Live Instances to Credential Theft

GitGuardian researchers found 321 n8n instances accepting API tokens exposed in public GitHub commits and demonstrated four ways attackers could use them to access sensitive data and downstream credentials without exploiting a software vulnerability.

We scanned public GitHub commits for exposed n8n API tokens and identified 4,576 unique credentials associated with 1,255 hostnames. Of the 896 instances reachable at the time of testing, 321 accepted at least one leaked token.

That means leaked credentials provided authenticated access to 36% of the reachable instances we tested, or roughly 26% of all hostnames identified in the commits.

The implications extend well beyond n8n. Organizations use the automation platform to connect databases, source code repositories, cloud environments, artificial intelligence services, customer support platforms, and other internal systems. A sufficiently privileged n8n token can expose workflow definitions and execution data, allow attackers to use stored credentials, and, in some configurations, enable them to extract the underlying credential values.

To measure the potential blast radius, we reproduced four practical attack techniques in a controlled n8n environment. Each required only documented REST API functionality and standard HTTP requests. No CVE exploitation or specialized tooling was necessary.

Why n8n is a high-value target

n8n is an open-source, low-code workflow automation platform with AI agent support and hundreds of built-in integrations. Organizations use it to connect internal tools, automate pipelines, implement business logic, and orchestrate API integrations across their technology stacks.

The platform can be self-hosted or deployed through n8n.cloud, and its open-source repository has attracted nearly 200,000 GitHub stars.

An n8n instance runs workflows composed of nodes. Some nodes trigger workflows on a schedule or through webhooks, while others transform data, execute code, or connect to external services using stored credentials such as API keys, tokens, and database passwords.

Those credentials are encrypted at rest using a master secret called N8N_ENCRYPTION_KEY. But n8n still needs to decrypt and use them whenever a workflow runs. An attacker with sufficient API privileges may therefore be able to reference those credentials in new workflows and make the instance use them on the attacker's behalf.

With more than 100,000 instances visible through Shodan and more than 50 security advisories published since January 2026, n8n has attracted the same attention as other high-value integration platforms.

As of March 31, 2026, 58% of the instances we scanned were running a version affected by at least one known security advisory. Several recent CVEs allowed attackers to escape execution sandboxes and gain arbitrary read or write access to the host filesystem.

CVE-2025-68613, an expression injection vulnerability with a CVSS score of 9.9, was added to the U.S. Cybersecurity and Infrastructure Security Agency's Known Exploited Vulnerabilities catalog on March 11, 2026, confirming exploitation in the wild.

Leaked API tokens create a separate risk. An attacker does not necessarily need to exploit an n8n vulnerability if a valid credential already provides authenticated access to the instance.

We found 321 instances accepting leaked tokens

GitGuardian Public Monitoring scans public sources for exposed credentials. For this research, we collected every n8n API token it had identified in public GitHub commits since April 2025.

Our pipeline extracted the n8n hostname committed alongside each token, sent a read-only validation request to the associated instance, and recorded the response.

The scan produced:

Stage Count
Unique API tokens 4,576
GitHub commits containing tokens 5,469
Unique hostnames extracted 1,255
Publicly reachable instances 896
Instances accepting a leaked token 321

The 321 confirmed instances represent approximately 36% of the 896 reachable instances and 26% of all 1,255 hostnames identified in the commits.

We ran the same process against n8n Model Context Protocol API keys found in the same commit set. MCP tokens allow AI assistants to call n8n workflows through the Model Context Protocol, making them a newer exposure surface than the REST API.

Of 372 MCP tokens identified, seven were still valid at the time of testing, or roughly 2%.

Why leaked n8n tokens can remain valid

An n8n API key is a signed JSON Web Token with an "aud": "public-api" audience claim. A decoded token looks like this:

{
  "sub": "efdf9cca-049a-46aa-afdc-172f0824f6cb",
  "iss": "n8n",
  "aud": "public-api",
  "jti": "aac8a7a8-c8c4-4855-8e8b-2806e90b16e1",
  "iat": 1781551662
}

The token records its issuance time in the iat claim. Older n8n API keys frequently contain no exp claim defining when they expire.

n8n introduced a 30-day default expiration in version 1.78.0 in February 2025, but many of the tokens found during the research had been generated without an expiration date. A key committed to GitHub months earlier could therefore remain usable until someone explicitly deleted or revoked it.

In practice, n8n API keys behave differently from self-contained JWTs that can be validated using their signatures alone. The key must also still exist in the n8n database. A token exposed in GitHub remains dangerous as long as the instance continues to recognize it.

Testing a candidate token requires one read-only request with the key passed through the X-N8N-API-KEY header:

curl -s -o /dev/null -w "%{http_code}" \
-H "X-N8N-API-KEY: <token>" \
https://n8n.example.com/api/v1/workflows

GET /api/v1/workflows returns workflow definitions available to the authenticated user.

A 200 response confirms that the token is accepted. A 401 indicates that the token is invalid, has been removed from the database, or failed signature verification. A 404 can indicate that the public API is disabled on the instance.

The request makes no changes to the target instance.

The instance URL is often committed beside the token

An n8n API key is useful only when an attacker can identify the instance that accepts it. In public GitHub commits, however, the hostname and token frequently appear together.

A .env file is one common example:

N8N_URL="https://n8n.redacted.cloud:5678"
N8N_API_KEY="eyJhREDACTEDPWw4"

We also found a newer pattern associated with Claude Code permission files.

Claude Code can store permitted shell commands in .claude/settings.json or .claude/settings.local.json. When users configure Claude Code to interact with n8n, they may place both the instance URL and API key directly inside an approved curl command.

Bash(curl -s "https://automation.redacted.fr/api/v1/workflows" \
-H "X-N8N-API-KEY: eyJhREDACTEDegE")

These settings files can then be committed to a repository without the same .gitignore safeguards that developers commonly apply to .env files.

The same hostname-and-token pairing appeared under several other variable names, including:

N8N_MCP_URL
N8N_WEBHOOK_BASE_URL
process.env.N8N_URL
os.getenv("N8N_HOST", "...")

Because the hostname was usually available in the same commit as the token, our pipeline did not require a separate infrastructure discovery step.

What an authenticated n8n token exposes

An n8n API token provides access according to the permissions of the user who created it. In practice, many of the exposed tokens appeared to belong to instance owners or administrators, likely because those were the users configuring the integrations and committing the keys.

Depending on the account's role, the public REST API may expose:

  • GET /api/v1/users: Usernames, email addresses, account creation dates, and pending invitations. Some information is restricted to instance owners.
  • GET /api/v1/workflows: Full workflow definitions, including node configuration, JavaScript or Python code in Code nodes, SQL queries, and secrets hard-coded in workflow parameters.
  • GET /api/v1/credentials: Credential names, types, and sharing information, but not the underlying values. This endpoint is restricted to owners and administrators.
  • GET /api/v1/executions: Workflow execution history. Adding ?includeData=true may return the complete input and output payload from each run.
  • GET /api/v1/data-tables: Rows from tables visible to the authenticated user.
  • GET /api/v1/variables: Variable names and their contents. This endpoint is restricted to owners and administrators.

Workflow definitions create the most immediate exposure because the API returns complete node configurations. If a developer placed an API key or token directly in a node parameter rather than using n8n's credential store, the value may appear in plaintext.

The credential endpoint itself does not return stored secret values. However, as our controlled tests demonstrate, an attacker with permission to create and execute workflows may be able to reference a stored credential and make n8n use or transmit it.

The audit endpoint provides an attack map

n8n's audit endpoint can provide an authenticated user with a security report for the instance:

curl -H "X-N8N-API-KEY: $JWT" \
-d "{}" \
-H "Content-Type: application/json" \
https://$N8N_INSTANCE/api/v1/audit

The response may identify:

  • Potential SQL injection exposures in workflows
  • Nodes with filesystem access
  • Unprotected webhooks
  • The running n8n version, which can be matched against known CVEs
  • Unused credentials
  • High-risk or community-installed nodes
  • Enabled security features
  • Node allowlists and blocklists
  • Telemetry settings

For a legitimate administrator, this information supports security reviews. For an attacker holding a leaked privileged token, it can provide a prioritized map of the instance's most promising attack paths.

Four attack techniques against a fully patched instance

GitGuardian did not perform the following exploitation techniques against exposed third-party systems. We reproduced them in a controlled n8n deployment built specifically for the research.

The test workflow contained three deliberate weaknesses:

  • A web form stored submissions in a data table accessible through authenticated workflows.
  • An OpenAI node processed each submission using a stored credential object.
  • An HTTP Request node published the output to GitHub using a token hard-coded in its node parameters.

Using that environment, we demonstrated four techniques, progressing from passive enumeration to active credential exfiltration.

Example n8n workflow

Technique 1: Enumerating the instance

GET /api/v1/users returned four accounts: the instance owner, two active users, and one pending registration.

GET /api/v1/workflows returned nine complete workflow definitions. In the target workflow, the parameters of an HTTP Request node contained a GitHub token in plaintext.

This first technique required no workflow modification. The exposed information was already available through read operations permitted to the authenticated account.

Technique 2: Using a stored OpenAI credential

GET /api/v1/credentials listed every stored credential object, including one named "OpenAI account." The endpoint revealed its name, type, and identifier, but not the API key itself.

We created a workflow with a Schedule trigger and an OpenAI node referencing the credential by its ID, then activated it.

Two tricks make this work. The Schedule trigger automatically fires after roughly 10 seconds, giving the workflow time to complete. GET /api/v1/executions?includeData=true then retrieves the complete execution record. Because n8n persists every node's full output, the OpenAI response appears in plaintext.

We successfully ran arbitrary OpenAI prompts using the instance's stored credential without ever seeing its value.

Technique 3: Reading the data table

The same two tricks apply.

We created a workflow with a Schedule trigger and a Data Table node configured to retrieve all rows, then activated it. GET /api/v1/executions?includeData=true returned the execution record seconds later, with every row in plaintext.

Four rows were exfiltrated, including names, email addresses, form responses, and processing statuses.

The workflow was then deleted.

Technique 4: Exfiltrating the raw OpenAI credential

The fourth technique went beyond using a stored credential and extracted its underlying value.

We started an HTTP listener, then created a workflow with a Schedule trigger and an HTTP Request node.

The key trick is that the HTTP Request node can use a stored n8n credential as its authentication method while sending requests to any URL. We configured the node to use the stored OpenAI credential and pointed it at our listener.

When the workflow fired, n8n attached the credential value as a Bearer token in the outgoing Authorization header. The listener captured the raw API key seconds after activation.

The workflow was then deleted.

Together, these techniques show how an attacker can progress from a leaked n8n token to broader credential and data exposure using legitimate platform functionality:

  1. Enumerate users, workflows, and security configuration.
  2. Identify stored credential objects and hard-coded secrets.
  3. Use stored credentials without viewing their values.
  4. Read data available to workflows.
  5. Cause n8n to transmit a stored credential to attacker-controlled infrastructure.

Deleting the malicious workflow also removed the associated execution records from the interface, potentially leaving defenders with limited evidence to investigate.

Real-world workflows showed similar weaknesses

The controlled demonstration was not based on a purely theoretical configuration. During the research, we found real n8n instances containing similarly exposed patterns.

One workflow automatically backed up its own definitions to a public GitHub repository. An SSH deployment key had been hard-coded directly into one of its nodes.

The repository's Git history contained earlier versions of every workflow, and the SSH key remained valid.

The workflow was effectively publishing its own sensitive configuration and credentials each time it ran.

The case illustrates why workflow automation platforms can create unusually large blast radii. They sit between multiple systems, process sensitive data, and routinely authenticate to external services. A weakness in one workflow can expose access far beyond the automation platform itself.

Responsible disclosure produced limited responses

Finding a valid exposed credential is only the first step. The risk remains until the affected organization revokes it and addresses any downstream exposure.

We attempted responsible disclosure with seven organizations:

  • Three hosting providers collectively associated with approximately 100 affected instances
  • Four individual companies

One hosting provider did not respond. Three of the four individual companies also did not respond.

One company operated a bug bounty program, acknowledged the report, paid a $1,200 bounty, and revoked the credential immediately. That combination of recognition and rapid remediation was the exception.

GitGuardian also made several disclosures directly to n8n during the research. n8n acknowledged the reports, said it was aware of the issues and planned to address them, and subsequently closed the reports. At the time of publication, GitGuardian had not independently confirmed that the related fixes had been released.

Approximately 30% of the 321 affected instances were hosted on n8n.cloud or similar managed services.

GitGuardian Public Monitoring already identifies exposed n8n API tokens and notifies affected developers through the company's Good Samaritan disclosure program. The findings also led GitGuardian to update its n8n API key detector and validity checks to improve detection accuracy.

Takeaways

A leaked n8n token is not an isolated credential exposure. A sufficiently privileged token can expose workflow definitions, hard-coded secrets, execution data, and internal tables. It can also allow an attacker to use stored credentials or, by creating a workflow that sends them to an external endpoint, extract their underlying values.

No CVE or specialized tooling was required in our controlled tests. A few standard HTTP requests were enough to move from an exposed token to sensitive data and downstream credential access. An attacker could then delete the workflow and its associated execution records, leaving defenders with limited evidence inside n8n itself.

Revoking the exposed n8n token is the first step, but it may not be the last. Organizations should determine which workflows, data, and downstream credentials the account could access, review the instance for unauthorized changes, and rotate connected credentials where exposure cannot be ruled out.

Automation platforms create a particularly large blast radius because they sit at the center of an organization's integrations. A single token can provide a path toward source control, databases, cloud services, AI APIs, support platforms, and customer data. The risk is defined not only by the n8n instance, but by every system connected to it.

About the author: Guillaume Valadon is a Staff Cybersecurity Researcher at GitGuardian, the secrets visibility and intelligence platform for securing the credentials that let code, machines, and AI agents access systems and act as trusted identities. Because attackers do not need to break in when they can log in with a valid credential, GitGuardian finds the secrets that matter across the whole secrets surface, inside and outside the perimeter, whether they are vaulted, stored, or leaked. It reveals their context and blast radius, then drives remediation at scale before one credential becomes a breach path. Trusted by 600,000+ developers and enterprises, including Snowflake, ING, BASF, Datadog, Qlik, Euronext, and Orange.

Found this article interesting? This article is a contributed piece from one of our valued partners. Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.



from The Hacker News https://ift.tt/pUEj2Kk
via IFTTT

Open VSX Removes 77 Malicious Evil Twin Extensions Exfiltrating Developer Data

A cluster of 77 extensions on the Open VSX marketplace has been found to impersonate legitimate developer tools while transmitting information about the systems and development environments on which they were installed.

The "evil twin" extensions were uploaded to the repository between July 26 and August 1, 2026, according to Manifold Security. The packages have been removed from Open VSX as of August 3, 2026.

"In most of the packages it sends little more than the machine's hostname," security researchers Ax Sharma and Cody Nash said. "In nineteen of them it sends a detailed description of the machine, the repository open in the editor, and the CI system the editor is running inside."

Of the identified extensions, 58 have been described as lightweight tools designed to exfiltrate the hostname and, in some cases, the workspace folder name or editor version.

The rest are reconnaissance payloads that transmit the developer-related information: local hostname and operating system username, the editor's name, version, host kind and machine ID, the platform and architecture, the locale and timezone, and the open workspace's folder name and full file system path.

Both share the same data-exfiltration domain, as well as similarities in code and behavior. The names of the 19 extensions are listed below -

  • amd.gaia-vscode
  • artsy.artsy-studio-extension-pack
  • configcat.configcat-feature-flags
  • iotaledger.iota-move
  • marketplace.visualstudio
  • obyte.oscript-vscode-plugin
  • openeuphoria.vscode-euphoria
  • oss.sfmc-devtools-vscode
  • rumbledb.jsoniq-vscode
  • ssagov.uef-snippets
  • taskfile.vscode-task
  • doi.fileheadercomment
  • mengsiCode.vscode-django-boilerplate
  • move.move-analyzer
  • uavcan.dsdl
  • vs-publisher-988541.apexsql-power-tools
  • casualjim.gotemplate
  • jcamp.dotnet-test-provider-view
  • superposition.supertoml-analyzer

What's notable about the campaign is that it reuses the names, namespaces, and descriptions of real Open VSX extensions, but they are published through unrelated accounts and assigned a low version number (e.g., 0.0.1). The main change involves swapping the contents of the bundled "extension.js" file with capabilities to capture and transmit data, while framing the collection as "anonymous usage metrics."

None of the extensions offer the advertised functionality mentioned in their listings. Instead, they display a status bar item along with a message that states they are active, before firing the data exfiltration step. In all 77 extensions, the data is sent to "mangorbit[.]com," which was registered on July 15, 2026, 11 days before the first packages were published.

The extensions that are part of the second set cluster have also been found to carry out the following steps -

  • Inspect files in the workspace's .git directory to obtain Git remote hosts and organizations, the domain of the developer's configured email, the current branch, and the HEAD commit SHA hash.
  • Enumerate up to 60 installed extension IDs and pick up the proxy hostname from the environment
  • The names of any CI markers present, and separately the values of GITHUB_REPOSITORY, CI_PROJECT_PATH, the Azure DevOps collection URI, the Buildkite organisation slug, the CircleCI project username, the Codespace name and the Gitpod workspace context URL from CI environments
  • Read the editor's own telemetry opt-out setting, checks if it's enabled, and sends the status

Further analysis indicates that the malicious code has contingency plans in place to query a DNS TXT record to retrieve a fallback exfiltration URL in the event the primary domain is blocked or taken down. The recon variant also features a retry mechanism that triggers the collection later on, suggesting the objective goes beyond a simple one-shot experiment.

"In the recon variant, attempts come at roughly fifteen minutes, fifty minutes and three and a half hours, then every seven or eight hours, resuming on every editor restart and giving up only after seven days," the researchers explained. "A machine that is offline, firewalled, or behind a proxy that drops the first request gets asked again for a week."

"The recon variant checks whether the open workspace's own devcontainer.json or .vscode/extensions.json references the extension's ID, and reports the answer as a single flag. In plain terms, it distinguishes installs that a repository's configuration caused from installs a human chose. That is the field you would want if the question you were asking was how am I being pulled in, and by what."

The disclosure comes as 450 unique npm packages spanning 2,244 artifacts have been compromised as part of a new software supply chain attack to deliver an information stealer and leverage the stolen npm token to push trojanized versions containing the same malware. The compromise campaign has been codenamed ChainDrop.

"The malicious releases contain a Mini Shai-Hulud variant, a self-propagating credential-stealing worm delivered through a large, heavily obfuscated Bun-based JavaScript payload," Microsoft said. "The malware typically executes automatically through an npm preinstall lifecycle hook before package installation completes."

"The malware can also use stolen GitHub credentials to inject Claude and Visual Studio Code configuration files into repositories, establishing persistence and creating an additional developer-to-developer infection path."

Although the tradecraft resembles the tradecraft observed in past Shai-Hulud npm worm activity, the activity remains unattributed at this stage.

"This sample also shows techniques not documented in earlier Shai-Hulud reporting: it downloads a standalone Bun runtime to execute a bundled second stage, uses a modular dispatcher with separate GitHub and domain-based delivery channels, and plants autostart hooks in .claude and .vscode to reach developers and AI coding agents who clone the source," Socket said.

OX Security noted the ongoing supply chain attacks targeting npm calls for a security layer that goes beyond blocking install scripts and requiring two-factor authentication (2FA) for maintainer accounts.

"It needs granular permission control over what a package can and can't do," security researcher Moshe Siman Tov Bustan said. "npm packages should require permission before they can exfiltrate AWS keys and GitHub credentials in a single command."



from The Hacker News https://ift.tt/s9uU7GX
via IFTTT

CISA Flags Langflow RCE, Tomcat, and N-central Flaws as Actively Exploited

The U.S. Cybersecurity and Infrastructure Security Agency (CISA), on August 5, 2026, added three flaws to its Known Exploited Vulnerabilities (KEV) catalog, citing evidence of active exploitation in the wild.

The list of vulnerabilities is as follows -

  • CVE-2026-9198 (CVSS score: 9.8) - A code injection vulnerability in Langflow that allows unauthenticated attackers to achieve full remote code execution on default Langflow deployments. (Fixed in July 2026 with version 1.10.1)
  • CVE-2026-34486 (CVS score: 7.5) - A missing encryption of sensitive data vulnerability in Apache Tomcat that allows a bypass of EncryptInterceptor, a cluster component that adds pre-shared key encryption to messages sent between cluster nodes. (Fixed in April 2026 with versions 11.0.21, 10.1.54, and 9.0.117)

Also added to the KEV catalog is CVE-2026-18556 (CVSS score: 8.2), an authentication bypass vulnerability in N-able N-central. It's worth noting that an incomplete fix for this issue prompted N-able to issue a fresh patch, which is tracked as CVE-2026-18577 (CVSS score: 8.2).

While CVE-2026-18577 was placed in the KEV catalog on Monday, the latest development signals that both vulnerabilities are being exploited by threat actors.

There are currently no details on how the Langflow flaw is being exploited. However, security defects in the open-source artificial intelligence (AI) application development platform have been repeatedly weaponized by bad actors in recent months.

The exploitation of CVE-2026-34486, on the other hand, has been attributed to an AI-enabled autonomous hacking campaign orchestrated by a Chinese-speaking threat actor operating under the aliases knaithe and KnYuan. The threat actor, based in Zhuhai, China, is said to have leveraged DeepSeek, via the Hermes Agent framework, as an offensive operator to target internet-exposed devices.

When initial attempts to exploit a Langflow flaw (CVE-2026-33017, CVSS 9.8) breach failed due to the target environment's restrictive configurations, the AI agent is said to have conducted autonomous research to identify other higher-value vulnerabilities, including flaws in n8n, to find a way in.

Separately, the Chinese-speaking adversary has been found conducting manual operations using known vulnerabilities in Citrix NetScaler (CVE-2026-3055), Marimo (CVE-2026-39987), Apache Tomcat (CVE-2026-34486), and IKE VPN (CVE-2026-33824) endpoints.

"This actor attempted to exploit over 460 targets, leveraging a mix of autonomous and manual techniques," Palo Alto Networks Unit 42 said. "What's interesting is that the actor appeared to allow DeepSeek to narrow the targeting scope, likely to conserve AI compute."

"This autonomous process of target identification, sampling and narrowing of scope is notable because the system executed hundreds of hours of manual targeting analysis in mere minutes, while also managing its own compute resources."

Federal Civilian Executive Branch (FCEB) agencies have until August 7, 2026, to apply the necessary fixes and safeguard their networks from active threats.



from The Hacker News https://ift.tt/kR1Kvd9
via IFTTT

Tuesday, August 4, 2026

ChainDrop supply chain compromise: Anatomy of a self-propagating worm

Microsoft Threat Intelligence identified a large-scale npm supply chain attack affecting more than 400 packages across multiple unrelated publishers, including packages associated with major enterprise software ecosystems such as keyv, flat-cache, cache-manager, and others. The malicious releases contain a Mini Shai-Hulud variant, a self-propagating credential-stealing worm delivered through a large, heavily obfuscated Bun-based JavaScript payload. The malware typically executes automatically through an npm preinstall lifecycle hook before package installation completes.

Once executed, the malware searches developer workstations and continuous integration and continuous delivery (CI/CD) environments for npm, GitHub, cloud, and infrastructure credentials. It uses recovered identities to authenticate to npm, GitHub, Amazon Web Services (AWS), Kubernetes, and HashiCorp Vault, enabling it to enumerate packages, repositories, workflow secrets, cloud parameters, and secret-store values. Collected data is encrypted and transmitted through an attacker-controlled HTTPS endpoint, with GitHub repositories serving as a fallback exfiltration channel.

The payload’s most significant capability is automated propagation. After obtaining an npm publishing token, it enumerates packages available to the compromised identity, downloads their latest tarballs, inserts the malware and setup loader, adds a preinstall hook, increments the patch version, and republishes the modified packages. The malware can also use stolen GitHub credentials to inject Claude and Visual Studio Code configuration files into repositories, establishing persistence and creating an additional developer-to-developer infection path.

In this blog, we’re sharing our analysis of this supply chain attack, along with protection, detection, amd hunting guidance. Organizations that installed an affected package with lifecycle scripts enabled should treat the associated developer workstation or build runner as potentially compromised. Investigations should prioritize credentials accessible to the affected identity, unauthorized npm releases, unexpected repository or workflow modifications, suspicious cloud and secret-store access, and artifacts produced by affected build systems. Organizations should revoke and rotate exposed credentials from a known-clean environment and rebuild affected systems and downstream artifacts from trusted sources.

Attack chain overview

The campaign appeared as a rapid sequence of unauthorized patch releases across more than 400 npm packages maintained by otherwise unrelated publishers. Many malicious versions had no corresponding source-code commit, pull request, tag, or legitimate release, indicating that the attackers modified and published package tarballs directly rather than compromising each public source repository.

Affected releases typically added a preinstall lifecycle script that launched a malicious file, setup.mjs, contained within the package, which launched the large, obfuscated Bun JavaScript bundle included in the package. Because npm runs preinstall scripts before installation completes, the payload could execute on developer workstations and build runners before application tests or conventional security checks began.

After execution, the malware performs the following actions:

  1. Determines whether it is running on a developer workstation or in a CI/CD environment. On workstations, it detaches itself to continue after installation; on CI/CD systems, it remains in the active job to access workflow secrets, runner credentials, and OpenID Connect (OIDC) publishing permissions. Both paths could support further package or repository propagation when suitable credentials are found.
  2. Collects credentials from local files, environment variables, command-line tools, and GitHub Actions runner memory.
  3. Authenticates to npm, GitHub, AWS, Kubernetes, and HashiCorp Vault to enumerate additional accessible resources and secrets.
  4. Encrypts and exfiltrates collected data through an HTTPS channel, using GitHub repositories as a fallback.
  5. Uses recovered npm publishing access to modify and republish additional packages.
  6. Uses GitHub credentials to inject files into Claude and Visual Studio Code configurations across repository branches for persistence.

The payload’s  package-propagation routine downloads each publisher’s latest release, inserts itself, increments the patch version, and publishes the resulting archive. This mechanism can rapidly transform one compromised npm identity into many malicious package releases.

Figure 1. Attack chain.

0. Initial publisher access

Evidence points towards stolen maintainer credentials as the attack vector for initial compromise. Later propagation used stolen npm publishing tokens and, in targeted workflows, GitHub Actions OIDC publishing access.

1. Payload startup and background execution

 The malicious npm package uses a lifecycle hook to launch its bundle.

During preflight, the payload checks the environment, exits on Russian-language systems, avoids duplicate instances, and starts a detached copy in the background on developer systems.

Figure 2. Platform identification and execution.

In CI environments, the payload remains attached so it can access credentials available to the active build job.

2. Initial credential discovery

The payload first collects information that is immediately available from the local system, shell, and GitHub Actions runner.

Figure 3. Credential discovery.

The shell collector attempts to obtain the GitHub CLI token and captures the values of all process environment variables. The filesystem collector searches credential files, shell histories, cloud configuration, Secure Shell (SSH) keys, and other sensitive locations.

3. Cloud and secret store enumeration

The recovered code then creates dedicated collectors for cloud and infrastructure services.

Figure 4. Credential enumeration.

These modules do not merely scan files for token patterns; they use available credentials to call service APIs, verify access, and retrieve additional secrets permitted to those identities.

The following snippet shows the authentication attempt made using the found credentials:

Figure 5. Credential validation.

4. GitHub credential theft and enumeration

Discovered GitHub tokens are validated before being used for additional collection or repository access.

Figure 6. GitHub credential collector.

The payload checks token scopes, enumerates writable repositories, and identifies repositories where workflow execution could expose additional secrets.

6. GitHub Actions OIDC abuse

The payload also contains a targeted publishing path for GitHub Actions workflows configured as npm trusted publishers.

Figure 7. Re-publishing package using GitHub OIDC token.

Packages published through this route can carry valid provenance because the publication originates from a legitimate workflow identity.

7. Exfiltration and fallback

Collected results are serialized as JSON, gzip-compressed, and encrypted with a randomly generated AES-256-GCM using a randomly generated 32-byte key and 12-byte initialization vector (IV). The AES key is then encrypted with the attacker’s RSA public key using RSA-OAEP-SHA256.

The payload first attempts delivery through an attacker-controlled dynamic HTTPS endpoint. The active domain can change through on-chain contract (0xE1f2395ee43e45A1556EC6438a88c31B83493103, selector 0x53ed5143) or, as a fallback, from a cryptographically verified signed GitHub commit (Signed fallback marker: thebeautifulmarchoftime). If that channel is unavailable, it creates a public GitHub repository with the description Shai-Hulud: Here We Go Again.

Encrypted results are committed as files such as: results-<timestamp>-<counter>.json.

At the time of analysis, the live contract returns npm-cache[.]com. Earlier candidates include pypi-get[.]com and js-mirror[.]com.

Figure 8. Exfiltrating stolen information.

In one fallback path, a stolen GitHub token is added separately using double Base64 encoding. This token field is encoded, not encrypted.

8. Repository persistence and secondary spread

The payload can use stolen GitHub credentials to inject the malware and supporting setup files into eligible repository branches. The recovered code targets Claude and Visual Studio Code configuration paths, including .claude/settings.json, .claude/setup.mjs, .vscode/tasks.json, and .vscode/setup.mjs.

These changes create a secondary infection route: future Claude or Visual Studio Code activity can restart the payload even after the original npm installation has completed. In a conditional GitHub fallback path, the payload also attempts to install a token-monitor component that maintains credential access and contains a destructive handler if the monitored token is revoked.

Figure 9. Injecting the malicious code into development ecosystems.

9. Worm behavior: Package modification and publication

The npm tokens found in collected data are checked for package-write permission and two-factor authentication (2FA)-bypass capability.

Figure 10. Republishing the package using stolen NPM token.
Figure 11. Malicious update to existing package and republishing.

The propagation routine downloads a package’s latest tarball, copies the current malware bundle into it, adds a loader, and replaces its lifecycle scripts. This creates the worm-like propagation pattern: one stolen token can produce malicious patch releases across every package available to that publisher. This also explains why malicious releases frequently appeared as an otherwise ordinary patch-version increment without corresponding source commits or pull requests.

Mitigation and protection guidance

Microsoft recommends the following mitigations to reduce the impact of this threat.

  • Update npm CLI to npm CLI v11.10.0+  and use the npm CLI min-release-age feature.
  • Review dependency trees, lockfiles, artifact repositories, and CI caches for the five compromised versions, including transitive references.
  • Pin known-good package versions.
  • Purge npm and yarn caches on affected developer endpoints and build hosts, especially if the compromised tarballs were written into shared CI caches.
  • Rotate credentials and secrets from a clean host if a build system or workstation imported a compromised version, because second-stage execution can expose tokens and compromise build integrity.
  • Ensure that Microsoft Defender Antivirus cloud-delivered protection, Microsoft Defender for Endpoint telemetry, Microsoft Defender for Containers, and Microsoft Defender XDR investigation workflows are enabled across developer and CI assets.
  • Organizations that produce software artifacts should also review their own release hardening because this incident appears consistent with CI/CD pipeline abuse through GitHub Actions OIDC publishing. Defenders should review token scopes, workflow approvals, protected environments, release provenance, and anomaly detection around automated package publication. Supply chain response cannot stop at host triage; it must also include verification that the release process itself has not been subverted.
  • After remediation, validate recovery deliberately. Rebuild affected projects from a known-good dependency baseline, confirm that compromised hashes are absent from package caches and artifact stores, and review endpoint telemetry for any lingering NodeJS directory artifacts such as Math_Symbol.js, Math_init.js,  or names similar to math_<guid>.js, or suspicious node child processes. For development organizations that share base images or golden build runners, rebuild those images as well so future jobs do not silently inherit poisoned caches or post-compromise persistence.

Indicators of compromise (IOC)

IndicatorDescription
54dc7ea54a1317cca0e890a2770630cf7fa6c97813e0cb9d2caa93012b350668  setup.mjs (npm tarball preinstall loader)
fd3ca4007b225fdf8de7af4345a19179d5efa8c4bb9205f88cda806e5684b1eb  setup.mjs (.claude and .vscode repository loader)
9fc2570b7cef51c1b8df116d144d11ff4096357be7d2c4c6367cfc2509cf1bccMath_*.js
npm-cache[.]comC2 domain
pypi-get[.]comC2 domain
js-mirror[.]comC2 domain
hxxps[:]//npm-cache[.]com:443/routerC2 URL

Microsoft Defender XDR detections

Microsoft Defender XDR customers can refer to the list of applicable detections below. Microsoft Defender XDR coordinates detection, prevention, investigation, and response across endpoints, identities, email, and apps to provide integrated protection against attacks like the threat discussed in this blog.

TacticObserved activityMicrosoft Defender coverage
Initial access / ExecutionMalicious files embedded in compromised npm packages execute the embedded payload automatically through a malicious preinstall lifecycle hook.Microsoft Defender Antivirus
– Trojan:NPM/ShaiLoader.BY
– Trojan:NPM/MalBun.A
– Trojan:NPM/ShaiWorm.DAY!MTB

Microsoft Defender for Endpoint
– Suspicious Node.js process behavior
– Suspicious Node.js script execution
Execution / Defense evasionThe preinstall loader launches a heavily obfuscated Bun-based JavaScript payload designed to hinder analysis and evade Node.js-focused monitoring.Microsoft Defender Antivirus
– Behavior:Linux/SuspBunActivity.A
– Behavior:Win32/SuspBunActivity.A

Microsoft Defender for Endpoint 
– Suspicious usage of Bun runtime
– Suspicious installation of Bun runtime
– Suspicious Node.js process behavior
– Suspicious script execution via Bun
– Suspicious Node.js script execution  

Microsoft Defender for Cloud
– Suspicious npm supply-chain compromise activity detected
Credential access / CollectionThe malware searches developer workstations and CI/CD environments for npm, GitHub, cloud, Kubernetes, and secrets.Microsoft Defender for Endpoint
– Credential access attempt
– Suspicious cloud credential access
– Enumeration of files with sensitive data
– Suspicious access of sensitive files  

Microsoft Defender for Cloud
– Sha1-Hulud Campaign Detected: Possible command injection to exfiltrate credentials

Advanced hunting queries

Microsoft Defender XDR customers can run the following advanced hunting queries to find related activity in their networks:

Execution of the preinstall script

DeviceProcessEvents
    | where Timestamp > ago(3d)
    | where FileName in~ ("node", "node.exe")
    | where ProcessCommandLine in~ ("node setup.mjs", "node  setup.mjs")

CloudProcessEvents
    | where Timestamp > ago(3d)
    | where FileName in~ ("node", "node.exe")
    | where ProcessCommandLine in~ ("node setup.mjs", "node  setup.mjs")

Execution of second-stage JavaScript using Bun runtime

DeviceProcessEvents
    | where Timestamp > ago(3d)
    | where InitiatingProcessFileName in~ ("node", "node.exe")
    | where InitiatingProcessCommandLine in~ ("node setup.mjs", "node  setup.mjs")
    | where FileName in~ ("bun", "bun.exe")
    | where FolderPath contains "bun-dl-" or ProcessCommandLine has "node_modules"

Malicious JavaScript from malicious packages

DeviceFileEvents
| where Timestamp > ago(3d)
| where SHA256 in~ ("9fc2570b7cef51c1b8df116d144d11ff4096357be7d2c4c6367cfc2509cf1bcc", "fd3ca4007b225fdf8de7af4345a19179d5efa8c4bb9205f88cda806e5684b1eb", "54dc7ea54a1317cca0e890a2770630cf7fa6c97813e0cb9d2caa93012b350668")

Credential access by malicious JavaScript

DeviceProcessEvents
   | where Timestamp > ago(3d)
   | where ProcessCommandLine has_any ('gh auth token', 'gcloud config config-helper', 'az account get-access-token', "azd auth token")
   | where InitiatingProcessFileName in~ ("bun", "bun.exe")
   | where InitiatingProcessFolderPath contains "bun-dl-" or InitiatingProcessCommandLine has "node_modules"

Microsoft Security Copilot

Security Copilot customers can use the standalone experience to create their own prompts or run prebuilt promptbooks to automate investigation and response tasks related to this threat. Useful promptbooks for this activity include Incident investigation, Microsoft User analysis, Threat actor profile, Threat Intelligence 360 report based on MDTI intelligence, and Vulnerability impact assessment. Some promptbooks require access to Microsoft Defender XDR, Microsoft Sentinel, or related Microsoft security plugins.

For this campaign, Security Copilot can help analysts summarize affected devices, pivot from the package hashes to endpoint evidence, identify hosts that communicated with the IPFS path or C2 infrastructure, and build remediation actions such as cache purge, credential rotation, and containment sequencing for impacted developer systems and build runners.

Threat intelligence reports

Microsoft customers can use Microsoft Defender XDR Threat analytics and related Microsoft threat intelligence reporting to stay current on the malicious activity, indicators, detection coverage, and recommended response actions associated with this compromise. These reports provide investigation context, protection guidance, and updated intelligence that security teams can use to prevent, mitigate, or respond to related activity in customer environments.

As with other active supply-chain investigations, defenders should monitor for updated intelligence on package status, additional affected versions, infrastructure changes, and newly surfaced post-compromise tradecraft. Microsoft will continue to incorporate validated indicators and detections into Microsoft security products as the investigation evolves.

Learn more

For the latest security research from the Microsoft Threat Intelligence community, check out the Microsoft Threat Intelligence Blog.

To get notified about new publications and to join discussions on social media, follow us on LinkedInX (formerly Twitter), and Bluesky.

To hear stories and insights from the Microsoft Threat Intelligence community about the ever-evolving threat landscape, listen to the Microsoft Threat Intelligence podcast.

Review our documentation to learn more about our real-time protection capabilities and see how to enable them within your organization.   

The post ChainDrop supply chain compromise: Anatomy of a self-propagating worm appeared first on Microsoft Security Blog.



from Microsoft Security Blog https://ift.tt/xz8X2c3
via IFTTT

VMware Memory Tiering: A Complete Guide

The evolution of virtualization technologies continues to drive new approaches to hardware resource optimization. VMware vSphere Memory Tiering is one such innovation, enabling a significant increase in available memory capacity on ESXi hosts by leveraging ultra-fast NVMe storage devices.

In simple terms, Memory Tiering creates a two-tier memory architecture: hot (frequently accessed) data remains in conventional DRAM, while cold (infrequently accessed) data is offloaded to a dedicated NVMe device. This approach allows organizations to substantially expand the effective memory capacity of a server while maintaining acceptable performance levels. In today’s environment, where memory can account for up to 80% of a server’s cost, this technology can reduce infrastructure expenses by approximately 40% while increasing virtual machine density without requiring hardware platform upgrades.

 

VMware Memory Tiering architecture and its benefits

Figure 1. VMware Memory Tiering architecture and its benefits

 

Memory Tiering was first introduced as a tech preview in vSphere 8.0 Update 3 and became a fully supported, production-ready feature in VMware Cloud Foundation (VCF) 9.0, which is built on vSphere 9.0. The technology is designed for organizations seeking to increase memory capacity for their workloads with minimal additional investment. This guide explains how Memory Tiering works, its benefits and limitations, hardware and software requirements, deployment procedures, and available tuning options.

How VMware Memory Tiering works

Solution architecture

Memory Tiering is integrated directly into the ESX hypervisor and transparently implements a two-tier memory hierarchy for virtual machines. Each host gains an additional memory layer (Tier 1) based on a local NVMe device, operating alongside conventional DRAM (Tier 0). When Memory Tiering is enabled, the hypervisor automatically migrates inactive virtual machine memory pages from DRAM to NVMe storage, freeing high-speed memory resources for active workloads. When a virtual machine accesses offloaded data, the corresponding pages are seamlessly returned to DRAM.

A key design principle is that only virtual machine memory is tiered. ESX vmkernel memory and critical hypervisor data structures always remain resident in DRAM, ensuring consistent host performance and stability.

At its core, Memory Tiering resembles the traditional practice of swapping memory pages to disk, but it is implemented in a far more intelligent and efficient manner through the use of high-performance NVMe devices and deep integration with the ESXi memory scheduler. The solution leverages the Active Memory metric for each VM (see Figure 5) to identify which data is “cold” and can be temporarily placed on NVMe without a noticeable impact on performance. This approach makes it possible to significantly increase the total amount of memory available on a host while minimizing the effect on the “hot” data that remains resident in DRAM.

Benefits of Memory Tiering

The primary advantage of Memory Tiering is substantial memory cost reduction. Instead of populating servers with expensive additional DIMMs, administrators can deploy comparatively inexpensive NVMe SSDs and effectively double available memory capacity through a 1:1 DRAM-to-NVMe ratio. For example, a server equipped with 1 TB of DRAM can gain an additional 1 TB of tiered memory capacity using a dedicated NVMe device.

In workloads with relatively low active memory utilization, such as office-oriented VDI environments where only 5-10% of allocated memory is actively used, it may be possible to increase effective memory capacity by 300-400%. This is achievable as long as DRAM remains sufficient to accommodate the workload’s active working set.

Beyond cost savings, Memory Tiering improves infrastructure flexibility. Additional memory capacity can be introduced simply by installing NVMe devices and rebooting hosts, eliminating the need for server replacement or disruptive hardware upgrades. This allows organizations to scale memory resources quickly in response to workload growth.

 

Memory Tiering capacity planning model

Figure 2. Memory Tiering capacity planning model

 

Naturally, there are trade-offs. Memory pages stored on NVMe devices are accessed more slowly than pages residing in DRAM. While NVMe SSDs offer exceptional throughput and low latency compared to traditional storage media, they remain orders of magnitude slower than system memory. Consequently, accurate identification of cold memory pages is critical. If a workload’s active memory demand exceeds the available DRAM capacity, frequent page migrations between DRAM and NVMe may occur, resulting in increased latency and degraded application performance.

In the following sections, we will discuss workload suitability for Memory Tiering and the limitations that should be considered before deployment.

Requirements and compatibility

Before planning a Memory Tiering deployment, verify that your environment meets the following requirements:

  • VMware vSphere version. Memory Tiering requires the latest VMware platform release, specifically VMware Cloud Foundation 9.0. Both vCenter Server and all ESX hosts within the cluster must be upgraded to this version. The feature is considered production-ready beginning with VCF 9.0 and includes improvements in stability, security (including memory encryption support), and integration with vMotion operations.
  • Licensing. Memory Tiering is a VMware vSphere feature and requires appropriate licensing. It is typically included with enterprise-level VMware Cloud Foundation editions. Consult the licensing documentation for your specific edition to confirm feature availability.
  • Hardware requirements: NVMe. Each ESX host participating in Memory Tiering must have at least one available NVMe SSD dedicated to the feature. The device is used entirely or partially as Tier 1 memory, with support for up to 4 TB of tiered capacity per host. Important considerations: the device must be a locally attached PCIe NVMe SSD; supported form factors include U.2, U.3, M.2, EDSFF (E1.S, E3.S), and similar PCIe-based designs; external storage systems, including Fibre Channel, iSCSI, NFS, SATA SSDs, and SAS SSDs, are not supported. Tier 1 memory must reside on local, high-performance storage.
  • NVMe performance and endurance requirements. VMware specifies strict requirements for Memory Tiering devices due to the intensive read/write activity associated with memory paging. Recommended minimum specifications include:
    • Endurance: Class D or higher according to the vSAN classification system. This corresponds to an endurance rating of at least ~7,300 TBW and typically around 3 DWPD (three full drive writes per day) over the warranty period. Simply put, the drive must be designed to withstand very frequent write cycles.
    • Performance: Class F (100K–350K write IOPS) or Class G (more than 350K write IOPS). These performance levels are typically found in enterprise-grade NVMe SSDs, especially those based on PCIe 4.0 or PCIe 5.0 interfaces.

 

Recommended NVMe classes for Memory Tiering

Figure 3. Recommended NVMe classes for Memory Tiering

 

Many storage vendors market such drives as Mixed Use or Enterprise Mixed, indicating balanced, high-performance characteristics for both read and write workloads. As a practical rule of thumb, 3 DWPD or higher should be considered the minimum target when selecting NVMe devices for Memory Tiering.

  • NVMe device compatibility. It is recommended to verify the selected NVMe devices against the official VMware compatibility database, such as the Broadcom/VMware vSAN SSD Compatibility Guide. This helps ensure that the drive is supported by ESX drivers and meets the required endurance and performance classifications. Although vSAN itself is not required for Memory Tiering, VMware recommends using the vSAN Compatibility Guide when selecting NVMe devices because the workload characteristics and performance requirements are very similar. Cutting costs on Memory Tiering storage is not advisable – low-end consumer NVMe drives with limited endurance may fail prematurely under the sustained paging workload generated by the feature.
  • Supported NVMe form factors. The solution supports a variety of NVMe form factors, including standard 2.5-inch U.2/U.3 drives, modern E3.S modules, and compact M.2 devices – essentially any NVMe form factor that can be installed inside the server. The key requirement is that the device meets the performance and endurance criteria described above. For example, unused M.2 slots on the motherboard can be utilized if all primary drive bays are already occupied.

 

Selecting compatible NVMe devices

Figure 4. Selecting compatible NVMe devices

 

  • Additional requirements. Initial configuration requires access to the ESX shell or SSH because creating an NVMe partition for Memory Tiering is currently performed from the command line using ESXCLI or a PowerCLI script. Administrators should also plan for host reboots, as a restart is required for Memory Tiering to become operational after configuration is completed.

Workload compatibility and limitations

Even when the hardware prerequisites are met, it is essential to evaluate whether your workloads are suitable for Memory Tiering. The primary consideration is the proportion of actively used memory:

  • Active memory ≤ 50% of DRAM. Ideally, the combined active memory consumption of all virtual machines on a host should not exceed approximately 50% of the installed DRAM capacity. This recommendation stems from the default Memory Tiering configuration, which provides an additional amount of memory equal to the host’s DRAM capacity (a 1:1 DRAM-to-NVMe ratio). In other words, half of the total available memory remains high-performance DRAM, while the other half resides on the slower NVMe tier. If the active working set fits entirely within DRAM, application performance remains comparable to a traditional memory-only environment. However, if active memory requirements exceed available DRAM capacity, some frequently accessed pages may be forced to reside on the NVMe tier, resulting in increased latency and reduced performance. As a practical rule, keeping active memory utilization below 50% of physical DRAM provides a reasonable margin for maintaining predictable performance. Before enabling Memory Tiering, administrators should collect and analyze Active Memory statistics for their virtual machines.
  • Measuring Active Memory. VMware vCenter allows you to monitor this parameter. On the VM page, under Monitor > Performance (Advanced), you can display the Active Memory metric (using the Real-time view). Third-party tools are also commonly used. For example, RVTools collects active memory statistics for VMs (although it too provides only a point-in-time snapshot, so be mindful of workload peaks). It is recommended to measure memory activity during the busiest periods (overnight batch processing, peak business hours) and to account for possible fluctuations with sufficient headroom. Only then can you determine whether Memory Tiering is suitable for a particular cluster or group of VMs.

 

Monitoring Active Memory consumption

Figure 5. Monitoring Active Memory consumption

 

  • Workloads unsuitable for Memory Tiering. VMware explicitly identifies several workload categories that should not be placed on hosts with Memory Tiering enabled because of their sensitivity to memory latency or feature incompatibilities. These include:
    • Virtual machines requiring extremely low latency (real-time applications, high-frequency trading systems, and similar workloads).
    • VMs using secure isolated memory technologies such as SEV, TDX, or SGX. These technologies typically require all memory to reside in DRAM and often operate on encrypted memory regions.
    • Virtual machines protected by Fault Tolerance (FT). FT does not support Memory Tiering because it relies on reserved memory and synchronous replication of VM state.
    • Very large “monster VMs” with massive memory footprints (for example, 1 TB or more per VM). Their active working set is usually large as well, making it impractical from a performance perspective to page even part of their memory to NVMe.
    • VMs configured with Latency Sensitivity settings or using large memory pages, as these mechanisms fundamentally conflict with the concept of paging memory to secondary storage.

If such critical workloads are present, the best approach is to dedicate hosts to them with Memory Tiering disabled. Mixed environments can be configured more flexibly: the feature can be disabled on specific ESX hosts within a cluster or even selectively disabled for individual VMs (as discussed later in the advanced configuration section). Nevertheless, maintaining a uniform cluster configuration is preferable, as it simplifies management, DRS operation, and workload balancing.

Planning and deployment scenarios

Sizing: How much memory and what kind of NVMe drives are needed?

Proper NVMe sizing is critical to a successful Memory Tiering deployment. Two scenarios are possible:

  • Greenfield (new infrastructure): you are designing a VCF/vSphere cluster from scratch with Memory Tiering in mind. In this case, servers can be equipped with less DRAM and compensated with NVMe capacity. For example, if calculations indicate that each host requires 1 TB of memory, you could install only 512 GB of DRAM and add 512 GB of NVMe. The total effective memory capacity would still reach the required 1 TB, with half of it provided by less expensive flash storage. This approach lowers the cost of each node. Naturally, it assumes that the active portion of memory (approximately 256 GB) fits entirely within DRAM. Memory Tiering is a Day-2 operation, meaning it cannot be enabled during the initial deployment of a VMware Cloud Foundation cluster – there is no corresponding option in the deployment wizard. Therefore, the NVMe devices intended for tiering may initially be used by vSAN or remain unused. This topic is discussed later in the section on vSAN integration.
  • Brownfield (existing infrastructure): you already have a deployed vSphere or VCF cluster and simply want to add an NVMe memory tier. In this case, the process is straightforward: purchase and install the required NVMe devices and then configure Memory Tiering. If the cluster is already running vSAN, the new drives simply need to be detected by the system and must not be automatically claimed by vSAN. If vSAN is not present, there are no additional prerequisites – install the NVMe drives and proceed with configuration.

As a sizing rule, the NVMe device should be at least as large as the host’s DRAM capacity if you intend to use the default 1:1 memory expansion ratio. For example, a server with 768 GB of RAM requires at least a 768 GB NVMe drive to double the available memory to approximately 1.5 TB. Larger drives are perfectly acceptable and provide flexibility for increasing the memory ratio later. At present, the maximum supported Memory Tiering NVMe capacity is 4 TB per host. Even if a larger drive is installed, ESX creates a partition no larger than 4 TB and does not utilize capacity beyond this limit.

Most importantly, the DRAM-to-NVMe ratio is configurable. The default value of 100% (1:1) is intended as a universal setting suitable for most workloads. However, workloads with very low active memory requirements can benefit from more aggressive ratios. For example, a 1:2 ratio (Mem.TierNvmePct = 200) provides an additional 200% memory capacity via NVMe.

 

Configuring the memory expansion ratio

Figure 6. Configuring the memory expansion ratio

 

In other words, if a host contains 512 GB of DRAM and a 4 TB NVMe partition, a 1:2 ratio causes 1,024 GB of NVMe to be used as Tier 1 memory, resulting in approximately 1.5 TB of total memory.

A 1:4 ratio (400%) is the maximum supported configuration. With 512 GB of DRAM, ESX would use 2,048 GB of NVMe, increasing total memory capacity to roughly 2.5 TB. The entire 4 TB device would not be consumed, but multiple VMs with a combined memory footprint of about 2.5 TB could be accommodated. Naturally, such an aggressive configuration assumes that less than 20–25% of the total memory is actively used; otherwise, frequent accesses to the NVMe tier will begin to impact performance.

As you can see, having a large NVMe device (up to 4 TB) allows the amount of NVMe memory actually utilized to be adjusted simply by changing the ratio, without repartitioning the drive. This is extremely useful from a future-proof perspective: a larger NVMe device can be purchased upfront, a maximum-size partition (4 TB) created, and the environment initially operated at a 1:1 ratio. If workloads later become even “colder,” the ratio can be increased to 1:2 or 1:4, thereby utilizing more of the already installed NVMe capacity and expanding host memory without purchasing additional hardware.

However, changing the ratio is a serious step and should only be done after confirming that the new active memory requirements still fit within DRAM.

Example: Suppose a host contains 1 TB of DRAM, and measurements show that no more than approximately 300 GB is actively used. In that case, memory expansion via NVMe is perfectly reasonable. Adding a 1 TB NVMe device and using a 1:1 ratio results in 2 TB of total memory. Since the active 300 GB still resides comfortably within the 1 TB DRAM tier, performance remains unaffected. But why stop there? Installing a 2 TB or even 4 TB NVMe device provides additional flexibility. With a 2 TB NVMe device and a 1:1 ratio, only 1 TB of that device is used as tiered memory, so total memory reaches 2 TB (1 TB DRAM plus 1 TB NVMe) and the remaining 1 TB stays in reserve. The active 300 GB remains entirely within DRAM, leaving ample headroom. Increasing the ratio to 1:2 allows the full 2 TB of NVMe to be utilized, producing 3 TB of total memory capacity. A 1:4 ratio, however, is impossible in this example because a 2 TB NVMe device supports only a 1:2 expansion. Achieving a 1:4 ratio with 1 TB of DRAM would require at least a 4 TB NVMe device.

Thus, the maximum achievable memory expansion is constrained either by VM activity levels or by the maximum supported NVMe size. In the current release, that limit is 4 TB per host, which is more than sufficient for most environments, since VMware customers rarely deploy servers with more than 4 TB of DRAM per host.

Integration with vSAN and storage

Memory Tiering is often discussed alongside VMware vSAN because, at first glance, the concepts appear similar: vSAN (especially the original OSA architecture) also relies on fast and slow tiers – cache and capacity. However, there is no direct dependency between Memory Tiering and vSAN. They are independent technologies that can operate concurrently on the same hosts without sharing resources.

It is important to remember that an NVMe device used for Memory Tiering must be dedicated exclusively to that purpose. You cannot configure Memory Tiering on the same drive that hosts a vSAN datastore or local VM storage. Although it is technically possible in a lab environment to split a physical SSD into two partitions – one for vSAN and one for Memory Tiering – VMware strongly discourages this practice in production.

If the same drive is shared between vSAN and Memory Tiering, both subsystems will compete for the device’s I/O resources, negatively impacting both memory and storage performance. It is essentially the equivalent of saying, “The fuel tank is half empty, let’s fill the unused space with water” – clearly not a good idea.

Therefore, when designing a vSAN-based configuration, allocate a separate NVMe device for Memory Tiering on each host, outside of any vSAN disk group. If you are using a ReadyNode configuration with only two NVMe drives dedicated to vSAN ESA and no spare devices, an additional drive will be required.

During a greenfield VCF deployment, if the NVMe device intended for Memory Tiering is installed from the beginning, ensure that it is not automatically claimed during vSAN configuration. By default, when vSAN is enabled on a new cluster, all available drives may be automatically claimed into disk groups. To avoid this, either leave the Memory Tiering NVMe out of the server until vSAN configuration is complete (it can be hot-added later), or disable vSAN auto-claim and manually assign disks while excluding the designated NVMe device.

If vSAN has already claimed all NVMe drives, there is no need to panic. After deployment, the drive can be removed from the disk group (through the UI or ESXCLI), its partitions cleared, and then repurposed for Memory Tiering. These procedures are documented by VMware. Removing a drive from vSAN without data loss requires either replica redistribution or sufficient free capacity within the cluster.

For existing (brownfield) environments, if vSAN has not yet been enabled, simply configure vSAN while excluding the Memory Tiering devices as described above. If vSAN is already operational and new NVMe drives were purchased specifically for Memory Tiering, simply install them in the hosts. Since they were never part of any disk groups, vSAN does not affect them.

It is worth emphasizing that Memory Tiering coexists perfectly with vSAN on the same servers, provided separate devices are used. VMs can simultaneously reside on vSAN (virtual disks) while utilizing NVMe-backed memory (virtual memory). Even encryption features are independent: vSAN datastore encryption and Memory Tiering page encryption can both be enabled without conflict because they protect different data types using separate mechanisms.

If no free NVMe devices are available but Memory Tiering is still desired, one possible compromise is to repurpose existing drives. For example, you may remove one NVMe drive from the storage configuration (in vSAN OSA, from the cache/capacity tier; in ESA, by simply reducing capacity) and dedicate it to Memory Tiering.

To do this:

1. Verify that the drive meets the required specifications (Endurance Class D, Performance Class F, and so on).

2. If the drive belongs to vSAN, initiate a Remove Disk from vSAN operation and wait for data rebalancing to complete.

3. If it hosts a local datastore, migrate or remove all VMs and then delete the datastore.

4. Clear the partition table on the NVMe device so that ESX recognizes it as a blank device without existing file systems.

5. Proceed with creating a Memory Tiering partition and configuring the feature on the host.

To reiterate: do not use the same NVMe device for multiple purposes in production. In lab environments (nested virtualization or test setups), experimentation is possible, but the risks should be fully understood.

Deploying Memory Tiering in test labs

An interesting question is whether Memory Tiering can be evaluated in a nested environment (ESX running inside ESX). Officially, VMware does not support Memory Tiering at the outer layer in nested scenarios – that is, when the physical host is an ESX server without NVMe devices and you attempt to enable tiering for its nested ESX virtual machines. Such a configuration offers no benefit because the outer (physical) ESX host has no visibility into the memory activity of nested VMs.

However, Memory Tiering can be enabled inside virtual ESX hosts, i.e., at the inner hypervisor layer. This requires presenting an NVMe device to the virtual ESX instance. In vSphere, an NVMe virtual disk can be created and attached to the nested ESX host. Inside the guest ESX, the device appears as a standard NVMe drive, allowing partition creation and Memory Tiering configuration just as if it were physical hardware.

Of course, this virtual NVMe device is ultimately backed by a VMDK file on a conventional datastore, so no performance gains should be expected. Nevertheless, the nested approach is ideal for training and experimentation. You can go through the complete configuration process, observe how Memory Tiering effectively “doubles” memory on virtual hosts, and experiment with various settings without requiring expensive physical hardware. Naturally, such configurations are not suitable for production.

Configuring Memory Tiering step by step

After completing the prerequisite tasks – installing hardware, upgrading hosts, and analyzing workloads – you can proceed with Memory Tiering configuration. The process consists of several stages.

1. Preliminary checks

“Measure twice, cut once.” Before making any changes to hosts, verify that everything is ready and all design decisions have been finalized. Appropriate drives should be selected, compatibility confirmed, fault-tolerance requirements (including RAID, discussed later) determined, placement of critical VMs planned, and NVMe devices visible on all hosts without any existing partitions or datastores.

It is recommended to verify this manually in vCenter. Under Storage Devices on each host, the NVMe device should appear with the correct model, an Available status, and no VMFS or vSAN signatures.

If hardware RAID is planned for the NVMe devices (for example, a mirrored pair for redundancy), configure the RAID groups at the controller level beforehand. ESX should ultimately see a single logical NVMe volume, which will be used for Memory Tiering. VMware Cloud Foundation 9.0 supports NVMe RAID configurations, including Intel VROC and Tri-Mode HBAs, but these settings are configured outside of ESX.

Decide which hosts will have Memory Tiering enabled. It may be enabled across the entire cluster or only on selected hosts. Hybrid configurations are fully supported. In practice, most environments either enable it everywhere or nowhere for consistency. However, in an eight-node cluster, for example, two hosts carrying latency-sensitive workloads may remain non-tiered. Those hosts simply continue operating with their physical DRAM capacity. In such mixed environments, DRS should ideally be configured to prevent “special” VMs from migrating to tiering hosts and vice versa.

Once these considerations have been addressed, proceed to the core configuration steps.

2. Creating an NVMe partition

Why is a dedicated partition required? VMware intentionally separates NVMe space used as memory from all other uses. ESX creates a dedicated Memory Tiering Partition and formats it specifically for this purpose. At present, the partition must be created manually. Two methods are available:

6. ESXCLI on each host: esxcli system tierdevice create -d /vmfs/devices/disks/<device>. (Refer to the VMware documentation for the exact syntax, as it may vary between releases.)

7. PowerCLI or PowerShell. A script can connect to vCenter and iterate through all hosts. VMware’s official blog provides a sample script that discovers NVMe drives by model, displays the list, requests confirmation, and executes the equivalent ESXCLI command to create the partition.

Regardless of the method used, the host must be rebooted after creating the partition and enabling Memory Tiering, so plan a maintenance window accordingly.

In vSAN clusters, place hosts into Maintenance Mode (Ensure Accessibility) or use vSphere Configuration Profiles (discussed later) to automate rolling reboots.

Frequently asked questions:

  • Can I create Memory Tier partitions on two different drives in the same host for redundancy? Technically yes – ESX does not prohibit it – but there is no practical benefit. ESX can use only one NVMe device for memory tiering at a time. If multiple Memory Tier partitions are present, the host mounts only one of them during boot and may not always choose the same one. The second partition remains unused. No mirroring or aggregation occurs between them. Therefore, if redundancy is required, use RAID1 or another hardware-level solution instead of multiple partitions.
  • What if there is no free NVMe device available? Memory Tiering is impossible without NVMe. SATA SSDs or NFS storage cannot be substituted. NVMe support is hard-coded into the feature.

After creating the partition, verify that it exists. In vCenter → Host → Configure → Storage Devices, selecting the NVMe device should show a new partition of the expected size. Typically, it occupies the entire drive, or up to 4 TB, whichever is smaller. Any capacity beyond 4 TB remains unallocated.

3. Enabling Memory Tiering

The next step is to activate Memory Tiering at the ESXi level. Even if the NVMe partitions have already been created, the feature remains inactive until the hypervisor starts using them. Memory Tiering can be enabled in several ways:

  • Via vCenter: either at the cluster level or on individual hosts. Starting with VMware Cloud Foundation 9.0, Memory Tiering management is available directly from the UI – simply select the desired hosts and click Enable Memory Tiering. In VMware Cloud Foundation, this can be done after the initial deployment.

 

Applying Memory Tiering configuration with vSphere Configuration Profiles

Figure 7. Applying Memory Tiering configuration with vSphere Configuration Profiles

 

  • Via ESXCLI: using the appropriate command (for example, esxcli system settings kernel set -s MemoryTiering -v TRUE; consult the product documentation for the exact syntax).
  • Via PowerCLI: using a script similar to the one mentioned earlier, or through vSphere Configuration Profiles, the new vCenter mechanism that replaces Host Profiles and allows a consistent configuration to be applied across a group of hosts.

 

Memory Tiering status in ESX Advanced System Settings

Figure 8. Memory Tiering status in ESX Advanced System Settings

 

Enable the entire cluster at once or one host at a time?

Memory Tiering can be enabled on all hosts simultaneously or rolled out gradually. The preferred approach is to use vSphere Configuration Profiles. You define a profile with Memory Tiering = On, apply it to the cluster, and vCenter will automatically remediate the hosts sequentially, rebooting them one at a time while migrating VMs to the remaining operational hosts. This approach ensures that vSAN remains available throughout the process, as vCenter waits for data synchronization and object compliance before proceeding with the next host reboot.

Once Memory Tiering has been enabled and the host has restarted, additional information becomes available under Host > Configure > Memory. For example, the reported memory capacity increases by the expected amount, and the Monitor > Memory view exposes statistics related to Tier0 (DRAM) and Tier1 (NVMe) utilization.

4. Post-deployment validation

After all hosts have been rebooted and the feature has been enabled, verify the following:

  • Total memory capacity has increased on each host. For example, if a host originally had 256 GB of DRAM and a 256 GB NVMe memory tier, the reported memory capacity should now be approximately 512 GB.
  • The Mem.TierNvmePct advanced setting is present. By default, its value is 100 unless it has been modified. If memory tier encryption has also been enabled (see below), the corresponding setting should appear as well.
  • Run a test workload. For example, power on several VMs and allocate more memory than the host’s physical DRAM capacity. If the VMs continue operating normally, this confirms that part of their memory footprint is being served from the NVMe tier. In vCenter, the Memory column for such VMs may show consumption exceeding 100% of the host’s physical memory. Previously, this typically indicated swapping; with Memory Tiering enabled, it reflects NVMe-backed memory expansion instead.
  • Inspect memory statistics under Monitor > Memory. The host view displays how many pages reside in DRAM and how many are located on NVMe, along with tier-in and tier-out activity (page migrations between the two tiers). With proper sizing and workload planning, nearly all hot pages should remain in DRAM, and NVMe read activity should stay relatively low.

 

Monitoring Memory Tiering devices in vCenter

Figure 9. Monitoring Memory Tiering devices in vCenter

 

At this point, the basic configuration is complete and Memory Tiering is operational. The next section covers advanced capabilities and fine-tuning options that allow the behavior of the feature to be adapted to specific workload requirements.

Advanced Memory Tiering settings

By default, Memory Tiering is configured conservatively: a 1:1 ratio, page encryption disabled, and all VMs on the host participating in tiering. For most environments, this is sufficient. However, several advanced parameters are available to administrators for optimization purposes. These settings are configured either in the host’s Advanced Settings or in the settings of individual virtual machines.

Adjusting the DRAM:NVMe ratio

As mentioned earlier, Mem.TierNvmePct is the primary parameter that determines the amount of NVMe capacity used as part of the overall memory pool. It is located under Host > Manage > Advanced Settings.

 

Configuring the Memory Tiering ratio through the Mem.TierNvmePct advanced host setting

Figure 10. Configuring the Memory Tiering ratio through the Mem.TierNvmePct advanced host setting

 

The value is expressed as a percentage of the host’s DRAM capacity:

  • 100 (default) means NVMe capacity equals 100% of DRAM (a 1:1 ratio).
  • 200 means NVMe capacity equals 200% of DRAM (a 1:2 ratio, resulting in a total memory pool three times the original DRAM capacity).
  • 50 means NVMe capacity equals 50% of DRAM (a conservative configuration that adds only 50% more memory, effectively maintaining a 2:1 DRAM advantage).

Changing the ratio simply requires modifying this parameter on the host (or through the API/PowerCLI). However, VMware strongly recommends against changing the ratio without careful analysis. Before increasing Mem.TierNvmePct to 200, for example, you should ensure that the current active memory working set is significantly smaller than the amount of DRAM available (ideally no more than one third of the resulting total memory pool). Otherwise, DRAM may become saturated, leading to page thrashing – continuous page movement between the two tiers.

If a new value is applied, the change takes effect only after a host reboot, since the memory hierarchy must be rebuilt. In a production cluster, the ratio can be adjusted one host at a time using Maintenance Mode to avoid disrupting running workloads.

NVMe memory encryption

Traditionally, RAM has been considered volatile and therefore transient, making memory encryption unnecessary. However, the NVMe tier is persistent, and memory pages may remain on the device across reboots (although ESX attempts to sanitize them). For highly security-conscious environments, VMware introduced Memory Tiering encryption without requiring an external Key Management Server (KMS).

Two levels of encryption are available:

  • Host-level encryption: enabled at the host level, meaning that all pages offloaded to the NVMe tier are encrypted. The encryption key is generated by the hypervisor (AES-XTS, 256-bit) each time a VM or the host starts. The corresponding parameter is Mem.EncryptTierNvme = 1 (configured in the host’s Advanced Settings). By default, it is set to 0 (disabled). When set to 1, encryption is enabled for any data that the host writes to the NVMe tier. By enabling this mode, you protect the data of all VMs running on that host from unauthorized access in the event that the NVMe device is physically removed or its contents are accessed outside the server.

 

Enabling host-level Memory Tiering encryption using the Mem.EncryptTierNvme advanced setting

Figure 11. Enabling host-level Memory Tiering encryption using the Mem.EncryptTierNvme advanced setting

 

  • Per-VM encryption: a more granular option that allows encryption to be enabled only for selected VMs. For example, you may have 50 VMs, of which 5 are critical (containing personal or financial data). In that case, page encryption can be enabled for those VMs only, while leaving it disabled for the rest to avoid unnecessary encryption overhead. The setting is configured at the virtual machine level (in the VM’s advanced settings): sched.mem.EncryptTierNVMe = TRUE.

 

Configuring per-VM Memory Tiering encryption by setting sched.mem.EncryptTierNVMe = TRUE in the virtual machine’s advanced parameters

Figure 13. Configuring per-VM Memory Tiering encryption by setting sched.mem.EncryptTierNVMe = TRUE in the virtual machine’s advanced parameters

 

By default, this setting is absent (that is, it is effectively FALSE). When set to TRUE, the hypervisor encrypts only the pages of that particular VM when they are offloaded to the NVMe tier. The mechanism is identical to host-level encryption: the encryption key is generated when the VM powers on, remains valid only for the lifetime of that VM, and never leaves the host. During vMotion, the VM’s encrypted pages are first decrypted, transferred over the secure vMotion channel, and then re-encrypted on the destination host using a newly generated key.

It should be noted that host-level and per-VM encryption are not mutually exclusive – they simply overlap. If host-level encryption is enabled, all VMs are already encrypted, which makes the per-VM setting redundant. Conversely, if host-level encryption is disabled but a specific VM is marked for encryption, only that VM’s pages will be encrypted.

The encryption overhead is minimal: hardware AES-NI acceleration is used, and encryption operations are performed asynchronously during page movement. In practice, the performance impact is barely noticeable, although for extremely heavily loaded systems, encryption can be left disabled (the default setting).

Excluding individual VMs from Memory Tiering

What about workloads for which any migration of memory pages to NVMe is undesirable? Examples include in-memory databases (SAP HANA, Oracle SGA), real-time systems, and latency-sensitive transactional applications. Such workloads should remain entirely within DRAM. For these cases, VMware provides an advanced parameter that disables Memory Tiering for a specific VM: sched.mem.enableTiering = FALSE (in the VM’s settings).

 

Disabling Memory Tiering for a specific virtual machine by setting sched.mem.enableTiering = FALSE in the VM's advanced configuration parameters

Figure 13. Disabling Memory Tiering for a specific virtual machine by setting sched.mem.enableTiering = FALSE in the VM’s advanced configuration parameters

 

When this parameter is configured, the hypervisor keeps the entire memory footprint of that VM exclusively in Tier0 (DRAM) and never offloads any of its pages to NVMe, even if Memory Tiering is enabled for other VMs on the same host. In effect, Memory Tiering is disabled for that particular virtual machine.

This setting should be used for unsupported or latency-sensitive workload profiles, including: large VMs configured with Latency Sensitivity = High; Fault Tolerance (FT) virtual machines; services with strict latency requirements.

Setting the parameter to FALSE ensures that such VMs continue to behave exactly as they would on a conventional host, even when Memory Tiering is enabled. Keep in mind, however, that if the host’s DRAM becomes fully occupied by these excluded workloads, Memory Tiering itself provides no benefit: the NVMe tier remains unused, and no additional VMs can be consolidated because neither DRAM nor NVMe capacity is available for them.

NVMe redundancy and Fault Tolerance

When discussing Memory Tiering reliability, we mentioned one option: using hardware RAID for NVMe devices to avoid losing memory pages in the event of a drive failure. VMware Cloud Foundation 9.0 supports such configurations (RAID1, RAID5, and similar NVMe RAID layouts). However, the advantages and disadvantages should be carefully weighed:

  • RAID1 (mirroring) with two NVMe drives provides protection against the failure of a single device and introduces minimal latency, since writes are performed in parallel. However, it requires twice as many drives and possibly a RAID controller (unless VMD/VROC is used), which increases the overall cost of the solution.
  • RAID5/RAID6 offer protection with better capacity efficiency, but their controllers are typically optimized for larger block sizes, and small-page write performance may degrade significantly. In addition, they require a controller and multiple drives, making the solution more complex and expensive.
  • Without RAID: if an NVMe device fails, the “cold” pages residing on it are lost. What happens next? In practice, ESX does not crash. From the hypervisor’s perspective, the loss of an NVMe tier resembles a hardware memory failure. Affected VMs may suddenly crash (resulting in a guest operating system crash, such as a BSOD or a kernel panic), while other VMs continue running and the host itself remains online. VMware HA then restarts the failed VMs on other nodes. In other words, an NVMe failure leads to performance degradation and the loss of some VMs, but not the failure of the entire host. If you are willing to accept this risk (the probability of enterprise NVMe failure is relatively low), you can forgo RAID altogether and rely on the standard HA cluster to recover affected workloads.

Ultimately, the choice between the cost and complexity of RAID and the risk associated with running without it is yours. One practical consideration is that if you are using vSAN ESA, cluster nodes typically do not have hardware RAID controllers at all, since ESA requires direct NVMe access through HBA controllers. In such environments, implementing RAID solely for Memory Tiering can be difficult. On more traditional servers equipped with RAID controllers, a two-drive NVMe mirror can be configured using Intel VROC, provided the CPU supports VMD.

Conclusion

VMware Memory Tiering represents a significant step forward in memory management for virtualized environments. Similar to tiered storage architectures, memory itself has become hybrid – a combination of fast but capacity-limited DRAM and slower but more cost-effective NVMe flash. When applied correctly, this technology allows organizations to substantially increase VM density, reduce the cost of new deployments by purchasing less DRAM (which is especially important today), and gain greater flexibility when scaling infrastructure, since adding memory capacity becomes faster and less expensive than before.

However, Memory Tiering is a sophisticated tool. It requires architects to carefully analyze workload characteristics and verify that the active memory footprint fits within physical DRAM. Appropriate hardware – enterprise-class NVMe SSDs – must be selected, and application-specific considerations must be taken into account, as not every workload benefits from running on tiered hosts. Without proper planning, Memory Tiering may fail to deliver the expected results or even cause performance degradation – for example, when attempting to tier “hot” database workloads or when overcommitting hosts beyond reasonable limits.

On the other hand, when deployed according to best practices, Memory Tiering operates almost transparently. Thousands of pages migrate between DRAM and NVMe without administrator intervention, while users continue to run their applications as before, only with access to significantly larger memory capacity. Combined with other vSphere capabilities such as DRS, vMotion, and vSAN, this technology enables the creation of more cost-efficient, scalable, and resilient private cloud infrastructures.

FAQ

Is VMware Memory Tiering supported in production?

Yes. Memory Tiering became a production-ready feature in VMware Cloud Foundation 9.0, built on vSphere 9.0. The earlier implementation in vSphere 8.0 Update 3 was a technical preview.

How much memory can VMware Memory Tiering add?

The default 1:1 DRAM-to-NVMe ratio doubles the host’s effective memory capacity. Configurable ratios up to 1:4 are supported, with a maximum NVMe tier size of 4 TB per host.

Can the same NVMe drive be used for Memory Tiering and vSAN?

No. The NVMe device should be dedicated to Memory Tiering. Sharing it with vSAN or local VM storage creates I/O contention and is unsupported for production deployments.

Which workloads benefit most from Memory Tiering?

It works best for workloads with large allocated memory but relatively small active working sets, such as VDI environments. Latency-sensitive applications, in-memory databases, Fault Tolerance VMs, and very large VMs are poor candidates.

What happens if a Memory Tiering NVMe drive fails?

VMs with memory pages stored on the failed device may crash, but the ESX host can remain operational. VMware HA can restart the affected VMs on another host. Hardware RAID can reduce this risk when the server configuration supports it.



from StarWind Blog https://ift.tt/XgUAZiN
via IFTTT