Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Agentic Safety 危害分類參考

本頁是 harm_taxonomy.yaml 的不可執行繁中參考。來源只提供示範用起始集合,分類、預設嚴重度與行動不得直接套用於正式平台。法律、危機處置、兒少安全與申訴程序須另行審查。

嚴重度

ID來源定義摘要
info已觀察到但不構成違規,僅作為脈絡記錄
low輕微違規,可考慮標示或降低觸及等柔性措施
medium明確違規,可移除內容或限制使用者
high嚴重違規,可停權或禁止進入特定空間
critical迫近傷害或法律風險,來源預設自動處置並升級通報

critical 的來源描述不構成本地自動處置授權。正式系統必須先定義緊急事件、人員輪值、證據保存、必要通報、錯誤回復與申訴流程。

分流佇列

  • auto_action 用於來源所稱高信心、低歧義情境,但仍須經平台核准才可啟用
  • needs_review 只提出建議,由人工審查者決定
  • needs_classification 用於無法清楚對應分類的模式,避免強迫套用標籤

分類表

ID說明預設嚴重度預設行動預設佇列
spam大量未經請求的宣傳、詐騙連結或重複張貼lowcontent_removecontent_demotecontent_hideauto_action
coordinated_inauthentic_behavior多個帳號協同行動,操弄能見度或公共討論highcluster_flagaccount_suspendcontent_removeneeds_review
harassment針對個人或群體的辱罵、威脅或圍攻mediumcontent_removetimeoutban_from_spaceneeds_review
hate_speech依受保護特徵攻擊他人highcontent_removecontent_label_attachaccount_suspendneeds_review
violent_threats對人員或地點提出可信的實體暴力威脅criticalcontent_removeaccount_suspendauto_action
sexual_content性內容,涉及未成年人時採更嚴格子分類highcontent_removecontent_hideaccount_suspendneeds_review
self_harm鼓吹或以露骨方式呈現自傷或自殺highcontent_hidecontent_label_attachneeds_review
impersonation帳號或內容冒充其他真實人物或實體mediumaccount_suspendcontent_removeneeds_review
platform_manipulation操弄觀看、投票、推升等互動或散布機制mediumcontent_demotecluster_flagaccount_suspendneeds_review
off_platform_harm平台內接觸進一步導向平台外傷害,例如性勒索或誘騙轉移criticalaccount_suspendban_from_spaceneeds_review

來源另定義 spam_ringastroturfingtargetedbulkadult_nuditycsam 等子分類。csam 在來源標記 mandatory_escalation: true,實際通報義務、保存方式與處置流程必須依適用司法管轄及組織政策核定。

契約與界線

  • Report.categoryharm_labels.category 應使用本分類檔的 ID,不能由執行中的 agent 臨時發明新分類
  • ModerationDecision.severity 可由預設嚴重度起算,但證據支持時才能調高或調低
  • ModerationDecision.action 應限制在分類的 default_actions,除非 playbook 提供可稽核的例外理由
  • 無法歸類時送往 needs_classification,新分類應透過版本審查提出
  • YouTube 與 Discord 範例僅用於說明,不能作為單一平台的完備政策