Skip to content

About Micans聽#212

Description

@for-just-we

馃摑 Overall Description

Hi, I recently read the latest Micans paper about Micro-Service program analysis but I don't see the code in the main branch. Is it available in main branch?

馃幆 Expected Behavior

A guide for setup Micans

馃悰 Current Behavior

I don't find the entrance to Micans

馃攧 Reproducible Example

No response

鈿欙笍 Tai-e Arguments

馃攳 Click here to see Tai-e Options
{{The content of 'output/options.yml' file}}
馃攳 Click here to see Tai-e Analysis Plan
{{The content of 'output/tai-e-plan.yml' file}}

馃摐 Tai-e Log

馃攳 Click here to see Tai-e Log
{{The content of 'output/tai-e.log' file}}

鈩癸笍 Additional Information

No response

Activity

  1. zhangt2333 commented on Nov 17, 2025

    @zhangt2333
    Member

    Thank you for your interest in MICANS. Currently, the refactoring of MICANS has not been completed (and has lower priority than our new frontend (#206)), so the code has not been pushed to the latest Tai-e. However, I believe the following resources will be helpful to you.

    We have prepared an artifact to reproduce all experimental data presented in MICANS' paper. This artifact includes detailed documentation, as well as the source code of MICANS. The artifact can be downloaded from the following URL: https://doi.org/10.5281/zenodo.14043593.

  2. self-assigned this
    on Nov 17, 2025
  3. for-just-we commented on Nov 24, 2025

    @for-just-we
    Author

    Thank you for your interest in MICANS. Currently, the refactoring of MICANS has not been completed (and has lower priority than our new frontend (#206)), so the code has not been pushed to the latest Tai-e. However, I believe the following resources will be helpful to you.

    We have prepared an artifact to reproduce all experimental data presented in MICANS' paper. This artifact includes detailed documentation, as well as the source code of MICANS. The artifact can be downloaded from the following URL: https://doi.org/10.5281/zenodo.14043593.

    Thanks, great work. I notice in the paper you evaluate injection attacks and data leak detection with Micans. But I don't see the checkers in the codebase. Are the taint checkers integrated into Micans?

  4. zhangt2333 commented on Nov 24, 2025

    @zhangt2333
    Member

    Thanks, great work. I notice in the paper you evaluate injection attacks and data leak detection with Micans. But I don't see the checkers in the codebase. Are the taint checkers integrated into Micans?

    When you say 'taint checkers,' are you referring to Taint Configuration?

  5. YunFy26 commented on Jan 27, 2026

    @YunFy26

    @zhangt2333 Hello锛孖 noticed that the directory results/dyer/* contains multiple output files, such as call-edges-1.txt, call-edges-2.txt, ..., and methods-1.txt, methods-2.txt.

    When calculating the Recall ($R$), it is necessary to find the intersection between the static analysis results ($S$) and the dynamic analysis results ($D$). I would like to clarify the correct aggregation logic for these multiple dynamic traces. Specifically, should the static results $S$ be intersected with each individual $D_i$ and then summed, or should they be averaged?

    Based on standard program analysis metrics, I assume the "Ground Truth" should be the union of all dynamic traces. Could you please confirm if the following formula is the intended way to calculate Recall for this dataset?

    1. Aggregate the Dynamic Ground Truth ($D_{total}$):

    $$D_{total} = \bigcup_{i=1}^{n} D_i$$

    1. Calculate Recall ($R$):

    $$R = \frac{|S \cap D_{total}|}{|D_{total}|}$$

  6. zhangt2333 commented on Jan 27, 2026

    @zhangt2333
    Member

    Hello @YunFy26,

    Thank you for your interest in Micans and for your question regarding the aggregation logic for multiple dynamic traces.

    Each test driver triggers a series of dynamic behaviors, which are recorded and dumped into files named call-edges-<number>.txt and methods-<number>.txt.

    If we were to merge files with different <number> suffixes, the resulting dynamic behavior would be inaccurate. For example, merging call chains a1->b->c and a2->b->d would produce spurious call chains a1->b->d and a2->b->c, leading to an inflated Recall value.

    Therefore, we calculate Recall separately for each file and then compute the average across all files.

    FYI, the source code for the Recall calculation is located in /recall-calculator inside the micans Docker image.

  7. YunFy26 commented on Jan 27, 2026

    @YunFy26

    Hello, @zhangt2333

    If we were to merge files with different suffixes, the resulting dynamic behavior would be inaccurate. For example, merging call chains a1->b->c and a2->b->d would produce spurious call chains a1->b->d and a2->b->c, leading to an inflated Recall value.

    I鈥檓 not entirely sure if I鈥檝e captured all the nuances correctly. Based on my reading, here is how I understand the procedure.

    In the dyer/basemall dataset, each test driver triggers different execution logic. For instance:

    calledges-1.txt contains no calls from the package com.medusa.gruul.account. This implies the specific test driver for this run simply did not trigger behaviors in that module.

    In contrast, calledges-25.txt does contain calls such as:

    <com.medusa.gruul.account.service.impl.MiniAccountExtendsServiceImpl: com.medusa.gruul.account.api.entity.MiniAccountExtends findByShopUserId(java.lang.String)>/com.baomidou.mybatisplus.core.conditions.query.QueryWrapper.<init>/0   <com.baomidou.mybatisplus.core.conditions.query.QueryWrapper: void <init>()>
    

    This confirms that the edge is a true dynamic behavior of the system.

    The Problem with Averaging: If we calculate the Recall for each file individually and then take the average, the result may be statistically unreasonable. A static analysis tool is designed to find all potential valid paths. If an edge exists in calledges-25.txt but not in calledges-1.txt, a static tool that correctly identifies this edge should not be "penalized" by an average that includes runs where the edge wasn't triggered.

    I have a quick question regarding the recall rate calculation. I'm performing some comparative tests and want to ensure I'm not misinterpreting the evaluation metrics. This is by no means a critique of MICANS' performance; I just want to be as rigorous as possible in my replication.For example, if I perform a static analysis and output the final result as call-edges.txt, the recall rate is calculated as follows:

    I compare call-edges.txt with calledges-1.txt, extract the identical lines, and add them to the file final-edges.txt. Then I perform the same operation on the remaining calledges-x.txt files, adding all identical lines to final-edges.txt. Finally, I remove duplicate lines from final-edges.txt. The resulting set of edges represents the recalled edges.

    Do you think the above approach is reasonable? Or have I misunderstood something in calledges.-<number>.txt or elsewhere? I am sorry for the interruption to your busy schedule. Your insights would be extremely valuable to me, and I look forward to your response.

  8. zhangt2333 commented on Jan 27, 2026

    @zhangt2333
    Member

    @YunFy26

    A static tool that correctly identifies this edge should not be "penalized" by an average that includes runs where the edge wasn't triggered.

    This situation will not occur. The recall-calculator computes the call graph from endpointi on both the dynamic graphi and the static graph, then compares the two graphs to obtain recalli. A behavior that is not triggered by endpointi in dynamic graphi (i.e., call-edges-<i>.txt and methods-<i>.txt) but is triggered by endpointj in dynamic graphj will not be counted when calculating the recall for endpointi.

    FYI, the source code of recall-calculator is located in /recall-calculator inside the micans Docker image.

    I compare call-edges.txt with calledges-1.txt, ... The resulting set of edges represents the recalled edges.

    This approach is correct. Edges that appear in both the static and dynamic graphs are the real (a.k.a., recalled) edges identified by static analysis. This is the traditional way of calculating the recall metric鈥攊f there's a hit, it counts as a success. However, recall-calculator adds an additional constraint: an edge is only considered a hit if it is triggered from the endpoint.


  9. YunFy26 commented on Jan 27, 2026

    @YunFy26

    @zhangt2333 I am very grateful for your explanation. I've got it now. Thank you so much for your patience and for your valuable time.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions