範例政策
每項政策位於獨立資料夾,資料夾內的 policy.md 包含提供給 gpt-oss-safeguard 的 prompt 文字。
此結構沿用 openai/gpt-oss-safeguard/example_policies/spam 的單一政策配置,但改以 Markdown 儲存 prompt 資產。
對應的驗證資料集位於上游 ../datasets/,依相同 policy slug 命名為各政策專用 CSV 檔案。
初始政策如下。
- harmful-body-ideals
- graphic-violent-content
- graphic-sexual-content
- dangerous-content
- age-restricted-goods-and-services
- dangerous-roleplay
這些政策是 prompt 與分類 label 範例。翻譯後的語意、模型表現及誤判風險仍須以繁中資料集另行驗證。