Repository navigation
Update llamafirewall_baseline.py - #165
Open
MaddipatlaChetan24 wants to merge 1 commit into
Open
MaddipatlaChetan24 wants to merge 1 commit into
MaddipatlaChetan24 wants to merge 1 commit into
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What type of PR is this? (check all applicable)
Related issue: N/A
What changed?
In LlamaFirewallBaseline.analyze_conversation(), benign-case confidence_score scaling read max_confidence, a variable only ever updated inside the if is_threat: branch. For any benign task that variable was always 0.0, so the “scale scanner concern into 0.1–0.3 risk” branch was dead code — every benign task was reported with a flat confidence_score = 0.1, regardless of how close any scan came to the block threshold. Added a separate max_scanner_confidence tracker, updated on every non-threat detection, and used it for the benign-case scaling instead.
Why?
This silently discarded real scanner signal from the reported malicious-risk metric for every benign task, which feeds the paper’s confidence/cost figures.
How did you test it?
python3 -m py_compile on the changed file; traced the control flow by hand to confirm max_confidence is unreachable in the benign path before the fix, and that max_scanner_confidence is now populated from every non-threat scan.
Potential risks
Low — only affects the numeric confidence_score for benign results; detection logic (is_malicious) and threat-path confidence are unchanged.