Skip to content

[analyzer] Collect analysis time statistics per translation unit (#718) - #5077

Closed
mmido6039 wants to merge 1 commit into
Ericsson:masterfrom
mmido6039:feat-718-analysis-time-statistics
Closed

mmido6039 wants to merge 1 commit into
Ericsson:masterfrom
mmido6039:feat-718-analysis-time-statistics

Conversation

@mmido6039

Copy link
Copy Markdown
Contributor

Fixes #718

Problem

There is no way to tell how the analysis time is distributed. If an analysis is
slow, the user cannot find out which analyzer or which translation units are
responsible for it, so there is nothing concrete to optimize.

What was missing

The issue asks for five things, and two of them are already available today:

  • The whole analysis time is already reported (Analysis length: X sec) and
    stored in metadata.json under timestamps.
  • Re-running the analysis on a single translation unit is already possible
    with CodeChecker analyze --file <path>.

The remaining three - time per analyzer, average per translation unit, and the
slowest translation units - all need the same missing piece: check() in
analysis_manager.py analyzes one translation unit per worker process, but it
never measures how long that takes. Its result tuple carries the return code,
the result file and the source file, but no timing, so nothing downstream can
aggregate it.

Fix

Measure the wall clock time of every translation unit in check() and return it
with the rest of the result. worker_result_handler() groups these per analyzer
and stores the summary in metadata.json under analyzer_statistics:

"duration": {
  "total": 7.523,
  "min": 0.912,
  "max": 5.204,
  "avg": 2.508,
  "slowest": [ { "file": "/path/to/file.c", "duration": 5.204 } ]
}

The analysis summary also prints one line per analyzer:

Analysis time of clangsa: total 12.30 sec, average 4.10 sec, longest 8.12 sec (parser.c)

Notes on the approach:

  • The statistics are not stored in the database. As suggested in the issue,
    they are generated by the analyze command and kept with the reports, so they
    can be inspected in the terminal without running a server. No server, schema or
    migration changes are needed - the new key is simply ignored by the store logic.
  • A monotonic clock is used, because the system time may be adjusted while the
    analysis runs.
  • The sum per analyzer is normally larger than the wall clock length of the
    analysis, since translation units are analyzed in parallel. This is documented.
  • The new summary line is added to skip_prefixes in the analyze_and_parse
    test, because the measured times differ from run to run. It removes only the
    new lines, so the existing expected outputs are unchanged.

Testing

  • New unit test analyzer/tests/unit/test_analysis_time_statistics.py: the
    summary calculation, the cap on the slowest translation units, aggregation per
    analyzer into the metadata, skipped files being left out, and analyzers that
    analyzed nothing.
  • New functional test test_analysis_time_statistics in
    tests/functional/analyze/test_analyze.py: runs a real analysis and checks
    both the printed line and the contents of metadata.json.
  • Verified fails-then-passes by removing only the aggregation while keeping the
    helper functions - the wiring tests then fail and the calculation tests still
    pass.
  • tests/unit and tests/functional/analyze show no new failures against a
    clean-tree baseline.
  • pycodestyle clean and pylint 10.00/10 on every changed file.

Out of scope / follow-ups

  • An exception raised inside the map_async callback (worker_result_handler)
    is swallowed and the analysis then waits on the pool's ~1 year timeout, so a
    bug there hangs the analysis instead of failing. Worth handling separately.
  • The skipped flag in the result tuple of check() is always False; the
    branch handling it in worker_result_handler looks vestigial.
  • Storing or displaying these statistics on the server/GUI side, and profiling
    the store command, are deliberately left out - the issue discussion argues
    against putting this in the database.

…csson#718)

The analysis time of the individual translation units was not measured, so
it was not possible to tell which files or which analyzer made an analysis
slow. Measure the time spent on every translation unit and summarize it per
analyzer in the metadata file and in the analysis summary.
@bruntib

bruntib commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Sorry, but I would close this PR, because it is being implemented in another PR: #5037
The number of open-source contributions fortunately increase, but it also means a burden on ticket management. We'll need to invent a procedure to set assignee properly.

@bruntib bruntib closed this Sep 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Performance statistics for the analysis

2 participants