Open weights improve inspectability, but they do not provide a reproducible supply chain by themselves. For production use I would want immutable weight and tokenizer digests, a signed model card, training/fine-tuning lineage, exact inference configuration, and an eval suite tied to the deployed artifact. Runtime controls still matter too: tool allowlists, retrieval-source policy, output validation, and traceable policy versions. Trust becomes measurable when the deployed binary, configuration, and evaluation evidence can all be reproduced.
Ahmet Özel
Open weights improve inspectability, but they do not provide a reproducible supply chain by themselves. For production use I would want immutable weight and tokenizer digests, a signed model card, training/fine-tuning lineage, exact inference configuration, and an eval suite tied to the deployed artifact. Runtime controls still matter too: tool allowlists, retrieval-source policy, output validation, and traceable policy versions. Trust becomes measurable when the deployed binary, configuration, and evaluation evidence can all be reproduced.