Skip to content
Merged
Changes from 4 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
80 changes: 80 additions & 0 deletions SECURITY.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
<!--
# Copyright 2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved.
#
# Redistribution and use in source and binary forms, with or without
# modification, are permitted provided that the following conditions
# are met:
# * Redistributions of source code must retain the above copyright
# notice, this list of conditions and the following disclaimer.
# * Redistributions in binary form must reproduce the above copyright
# notice, this list of conditions and the following disclaimer in the
# documentation and/or other materials provided with the distribution.
# * Neither the name of NVIDIA CORPORATION nor the names of its
# contributors may be used to endorse or promote products derived
# from this software without specific prior written permission.
#
# THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS ``AS IS'' AND ANY
# EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
# IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR
# PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR
# CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL,
# EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO,
# PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR
# PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY
# OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
# (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
# OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
-->

# Security Policy

## Reporting a Vulnerability

NVIDIA is dedicated to the security and trust of our software products and services, including all source code repositories managed through our organization.

To report a potential security vulnerability, please use one of the following channels:

1. **NVIDIA Vulnerability Disclosure Program** (preferred): https://www.nvidia.com/en-us/security/
2. **Web form:** [Security Vulnerability Submission Form](https://www.nvidia.com/object/submit-security-vulnerability.html)
3. **Email:** [NVIDIA PSIRT](mailto:psirt@nvidia.com). Please encrypt sensitive reports with NVIDIA's [PGP key](https://www.nvidia.com/en-us/security/pgp-key).
4. **GitHub Private Vulnerability Reporting (where enabled):** use the "Report a vulnerability" button on the Security tab of this repository.

**Do not open a public issue or pull request to report a vulnerability.**

Please include:

* Product or component name and version or branch
* Type of vulnerability
* Steps to reproduce
* Proof of concept, if available
* Potential impact and how it could be exploited

See https://www.nvidia.com/en-us/security/ for past NVIDIA Security Bulletins and Notices.

## Security Architecture and Context

**Project:** The Triton TensorRT-LLM Backend.

**Software type:** Software component (library, backend, client or tool) used as part of a Triton Inference Server deployment.

**Security boundaries:** The main security boundary is between this component and the data, models and configuration it is given, and between it and the server or application that hosts it.

**Repository Exposure Classification:** Public.

**Service Exposure Classification:** Deployment-dependent. Exposure depends on how the software is deployed and configured by the operator.

## Threat Model

1. **Untrusted input:** Requests, models, configuration or data supplied to this component may be malformed or malicious, and could cause crashes, memory errors or unintended behavior if not validated.
2. **Supply chain:** Source and build dependencies fetched at build or install time may be compromised, outdated or unpinned.
3. **Network exposure:** When deployed behind a network-facing server, endpoints may be reachable by untrusted clients. This component does not by itself provide authentication, authorization or encryption.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 security Private MPI requirement removed The new network guidance discusses client-facing endpoints but drops the requirement to keep the MPI coordination fabric private. Supported orchestrator deployments can span nodes, and an HTTP-facing gateway does not protect traffic between those nodes. Operators could secure inference endpoints while leaving that communication reachable by untrusted peers; please retain the private-fabric assumption.

How this was verified: The deployment documentation supports MPI-based orchestrator operation across nodes, while the new assumption names only a trusted environment or gateway.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

4. **Resource exhaustion:** Oversized or numerous requests may consume memory, compute or other resources and degrade availability.
5. **Information disclosure:** Logs, metrics and error messages may reveal sensitive data such as paths, identifiers or request content.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 security Tenant isolation warning removed If a deployment serves multiple authorized tenants, gateway authentication alone does not isolate inference state. The previous warning about shared state is gone, while the documented LoRA cache lets subsequent requests use a cached adapter by supplying only its lora_task_id. Restoring the single-tenant or tenant-isolation assumption would help operators avoid treating an authenticated shared deployment as sufficient protection.

How this was verified: The LoRA documentation says subsequent requests can use a cached adapter by supplying only its task ID.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!


## Critical Security Assumptions

* The component is deployed in a trusted environment or behind a gateway that provides authentication, authorization, TLS and rate limiting.
* Models, configuration and other inputs come from trusted sources.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 security Artifact integrity guidance removed The new trusted-source assumption no longer tells operators to check model artifacts before loading them. The documented workflow downloads models and tokenizers from external sources, and preprocessing loads tokenizers when the server initializes a model. A trusted source alone does not establish that the downloaded artifact is the intended one; please restore the operator's integrity-check responsibility.

How this was verified: Repository documentation shows external model downloads and tokenizer loading during Triton model initialization.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

* Dependencies and the build environment are kept up to date and obtained from trusted sources.
* Operators protect secrets, certificates and credentials, and restrict access to logs and metrics.
* Host operating system, driver and hardware security are the operator's responsibility.
Loading