If you believe your technology systems are free from Chinese AI influences, Cisco’s recent research suggests a reevaluation might be necessary. The U.S. government views AI from China as a potential national security threat, prompting many to avoid it. However, Cisco’s new findings reveal that such country-based labels may not accurately reflect an AI model’s origins or characteristics due to what they describe as ‘provenance entanglement’.
Understanding Provenance Entanglement
In a detailed blog post, Cisco explains that the simple labeling of AI models by country does not fully represent the complexities of their development. Using two AI fingerprinting methods, researchers uncovered that model weights and behaviors do not reliably indicate their publisher’s name or country. This discovery raises concerns as models might incorporate elements from other nations, regardless of their labeled origin.
The study suggests that AI models need something akin to a Software Bill of Materials (SBOM), though challenges exist. The dependencies that influence a model’s behavior are often buried within the learned weights, not accessible in a straightforward manifest. Thus, a model labeled as originating from the U.S. could still hold characteristics derived from a Chinese model, and vice versa.
Implications for Security and Transparency
Cisco emphasizes that while the origin label provides insight into the accountable developer and jurisdiction, it does not offer a full view of the model’s internal workings. Using their Model Provenance Kit and VAIL’s Behavioral Fingerprinting, the researchers examined relationships between Nemotron and Qwen models, finding significant lineage connections.
This research highlights the importance of understanding model lineage, particularly as it relates to potential vulnerabilities such as backdoors or biases. Organizations must recognize which downstream models could be affected by issues discovered in upstream models.
Recommendations for Future AI Use
For enterprises, Cisco advises considering the publisher’s identity as just one aspect of the evaluation process. A thorough review should include examining the model’s lineage, training dependencies, and behavioral analysis. Regulators and AI developers are urged to improve transparency and understanding of a model’s upstream influences to mitigate risks effectively.
In geopolitical terms, relying solely on the labeled country of origin can lead to misconceptions and overlooked risks. Cisco proposes a ‘model bill of materials’ approach to improve AI model adoption responsibly. This would involve documenting key components and relationships within models, providing clearer insights into their true origins and potential vulnerabilities.
In conclusion, while country labels retain some significance, they do not define a model’s technical heritage. As Cisco’s researchers state, ‘Models do not have passports. They have supply chains.’
