Skip to content

Update llamafirewall_baseline.py - #165

Open
MaddipatlaChetan24 wants to merge 1 commit into
uber:mainfrom
MaddipatlaChetan24:patch-6
Open

MaddipatlaChetan24 wants to merge 1 commit into
uber:mainfrom
MaddipatlaChetan24:patch-6

Conversation

@MaddipatlaChetan24

@MaddipatlaChetan24 MaddipatlaChetan24 commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

What type of PR is this? (check all applicable)

  • Refactor
  • Feature
  • Bug Fix
  • Optimization
  • Documentation Update

Related issue: N/A

What changed?
In LlamaFirewallBaseline.analyze_conversation(), benign-case confidence_score scaling read max_confidence, a variable only ever updated inside the if is_threat: branch. For any benign task that variable was always 0.0, so the “scale scanner concern into 0.1–0.3 risk” branch was dead code — every benign task was reported with a flat confidence_score = 0.1, regardless of how close any scan came to the block threshold. Added a separate max_scanner_confidence tracker, updated on every non-threat detection, and used it for the benign-case scaling instead.

Why?
This silently discarded real scanner signal from the reported malicious-risk metric for every benign task, which feeds the paper’s confidence/cost figures.

How did you test it?
python3 -m py_compile on the changed file; traced the control flow by hand to confirm max_confidence is unreachable in the benign path before the fix, and that max_scanner_confidence is now populated from every non-threat scan.

Potential risks
Low — only affects the numeric confidence_score for benign results; detection logic (is_malicious) and threat-path confidence are unchanged.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant