Security and Risk Notice
The platform's API supports a wide range of applications, such as question answering, writing, and dialogue. While using our API can create convenience for end users, it may also give rise to security issues. This document is intended to help customers understand the potential security issues that may arise when using the API. This document first introduces how to safely call the API as part of a product or service, then enumerates several specific issues to consider, provides general guidance on risks, and specifically provides further guidance on robustness and fairness.
I. Security Challenges of Open Machine Learning Systems
We define API security as follows: the condition of being free from physical, psychological, or social harm to persons, including but not limited to death, injury, illness, distress, misinformation, extreme behavior, property damage, or damage to the environment. Our notices and guidance on API security are based on special considerations for systems incorporating Machine Learning (ML) components that engage in high-bandwidth, open-ended interactions with humans (e.g., via natural language).
The robustness of ML components is limited. An ML component can only be expected to provide reasonable outputs when inputs are relatively similar to the given training data. Even if an ML system is deemed safe when operating under conditions similar to the training data, unexpected inputs from users can place the system into an unsafe state, and users often cannot tell which inputs will or will not lead to unsafe behavior. Open-ended ML systems that interact with individuals (for example, in question-answering applications) are also vulnerable to adversarial inputs from malicious users who deliberately attempt to place the system into an unintended state. Therefore, as a mitigation measure, developers using this platform should manually evaluate model outputs for each considered use case, produced across a representative range of inputs and some adversarial inputs.
ML components are biased. ML components reflect the values and biases present in the training data, as well as the values and biases of their developers. Systems using ML components (especially those interacting in an open-ended manner) may perpetuate or amplify these values. Security concerns arise when values embedded in ML systems are harmful to individuals, groups, or key institutions. For an ML component like the API trained on vast amounts of value-laden training data collected from public sources, the scale of the training data and complex social factors make it impossible to completely eliminate harmful values.
Open-ended systems carry significant hidden risks. Systems with high-frequency interactions with end users, such as natural language dialogue or question answering, can be used for virtually any purpose. This makes it impossible to exhaustively enumerate and mitigate all potential security risks in advance. Instead, we recommend adopting an approach focused on broadly considering categories and contexts of potential harm, continuously detecting and responding to harm events, and continually integrating new mitigations as needs become apparent.
Safety is a continuous consideration in developing ML systems. The safety characteristics of an ML system change each time an ML component is updated, such as by retraining them with new data, or by training new components from scratch with new architectures. Because ML is an active research area and new performance levels are frequently updated as research progresses, ML system designers should anticipate frequent updates to ML components and plan to execute continuous safety analyses.
II. Harms to Consider in Risk Analysis
We will illustrate potential harms (or pathways to harm) that may arise in systems involving the API. The following examples are not exhaustive, and not every category applies to every application scenario; use cases vary in their degree of openness and risk. In identifying potential harms, developers should consider the system to be developed based on the context of use, including those who use the system and those affected by it, and investigate representative sources of harm.
Providing false information. The system may provide users with false information regarding safety or health issues, such as giving an incorrect response to a user inquiring whether they are experiencing a medical emergency and should seek medical care. The intentional creation and dissemination of misleading information via the API is strictly prohibited.
Discrimination. The system may persuade users to believe things harmful to certain groups, such as using racist, sexist, or ableist language.
Harm to individuals. The system may generate results that could harm individual humans, such as encouraging self-destructive behavior (e.g., gambling, substance abuse, or self-harm) or damaging their self-esteem.
Incitement of violence. The system may persuade users to commit acts of violence against any other person or group.
Physical injury, property damage, or environmental destruction. In certain use cases, such as where a system using the API is connected to physical actuators capable of causing harm, the system is central to safety concerns, and unintended behavior in the API could lead to malfunctions resulting in physical damage.
III. The Importance of Robustness
"Robustness" refers to a system working reliably as planned and expected in specific environments. Developers using this platform should ensure that their applications possess the robustness required for safe use, and should ensure that this robustness is maintained over the long term.
Robustness is a challenge. Language models such as those included in the API are useful for a range of purposes, but may fail in unexpected ways due to factors such as limited world knowledge. These failures may be visible, such as generating irrelevant or clearly incorrect text, or invisible, such as failing to retrieve relevant results when using API-driven search. Risks associated with using the API vary widely across use cases, though general categories of robustness failures to consider include: generating text irrelevant to context (providing more context makes this less likely); generating inaccurate text due to gaps in the API's current knowledge; continuing to supply offensive context, etc.
Context is crucial. Developers should keep in mind that the API's output depends heavily on the context provided to the model. Providing additional context to the model (such as giving high-quality examples of expected behavior prior to a new input) makes it easier to guide the model's output in the desired direction.
Human oversight. Even with substantial efforts to enhance robustness, some failures may still occur. Therefore, API customers should encourage end users to carefully review the API's outputs before taking any action (e.g., disseminating these outputs).
Continuous testing. Despite good initial performance, the API may fail to achieve expected results, one cause being shifts in input distributions over time. In addition, the Large Model Open Platform may provide improved versions of the models over time; developers should ensure that these versions continue to perform well under specific environments.
IV. The Importance of Fairness
"Fairness" here refers to ensuring that the API neither degrades performance based on user demographics nor generates text biased against certain groups. API users should take reasonable steps to identify and reduce foreseeable harms related to demographic biases in the API.
Fairness in ML systems is extremely challenging. Because the API is trained on human data, our models exhibit various biases, including but not limited to biases related to gender, race, and religion. For example: the API is primarily trained on Chinese text and is best suited for classifying, searching, summarizing, or generating such text. By default, the API performs less well on inputs that differ from the distribution of data on which it was trained, including non-Chinese languages and specific Chinese dialects that are underrepresented in our training data. The Large Model Open Platform provides information on some of the biases we have discovered, although this analysis is not exhaustive; developers should consider fairness issues that may be particularly salient in their use scenarios, even if not discussed in our foundational analysis. Note that context is critical here: providing insufficient context to guide generation, or providing context related to sensitive topics, is more likely to yield offensive outputs.
Characterize fairness risks prior to deployment. Users should consider their customer base and the range of inputs for which they will use the API, and should evaluate API performance across a variety of potential inputs to identify situations where API performance may degrade.
Filtering tools can provide some assistance, but are not a panacea. The platform has enabled automated filtering tools to flag potentially sensitive outputs, and is actively working with customers to test and refine this tool. The purpose of filtering tools is to help developers mitigate the risk of offensive outputs, but they are not suitable for all applications. Developers should consider whether their use case requires the use of such technology, and if so, how to adapt these technologies to best fit their use case. It is important to note that these tools are not a panacea for eliminating all potentially offensive outputs—offensive outputs using otherwise "safe" words may still be generated.
