Skip to content

Commit 879aea4

Browse files
committed
Convert 68 await-free tokio tests to sync
1 parent b9d964c commit 879aea4

36 files changed

Lines changed: 762 additions & 354 deletions

‎CONTRIBUTING.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
# Contributing to BACO
22

3-
Welcome! BACO (Bug Analysis & Cross-reference Orchestrator) is a research-backed SAST scanner that augments static analysis with LLM-powered discovery across a 23-phase pipeline. Sponsored by [Regolo.AI](https://regolo.ai), this project integrates techniques from 20 academic papers to detect vulnerabilities with higher accuracy than traditional tools.
3+
Welcome! BACO (Bug Analysis & Cross-reference Orchestrator) is a research-backed SAST scanner that augments static analysis with LLM-powered discovery across a 23-phase pipeline. Sponsored by [Regolo.AI](https://regolo.ai), this project integrates techniques from 31 academic papers to detect vulnerabilities with higher accuracy than traditional tools.
44

55
## Getting Started
66

‎README.md‎

Lines changed: 22 additions & 19 deletions
Original file line numberDiff line numberDiff line change
@@ -33,16 +33,15 @@ cargo build --release
3333
# 2. Install semgrep (required for static analysis)
3434
pip install semgrep
3535

36-
# 3. Set your LLM API key
37-
export MISTRAL_API_KEY="your-key-here"
36+
# 3. Set your LLM API key (per-phase or generic fallback)
37+
export LLM_DISCOVERY_KEY="your-key-here" # or LLM_API_KEY for generic fallback
3838

39-
# 4. Run pre-flight checks
39+
# 4. Initialize config and run pre-flight checks
40+
./target/release/baco init /path/to/project
4041
./target/release/baco doctor
4142

42-
# 5. Configure and scan
43-
cp config.toml my-config.toml
44-
# Edit my-config.toml: set [project] path to your target code
45-
./target/release/baco scan --config my-config.toml
43+
# 5. Scan
44+
./target/release/baco scan --config config.toml
4645
```
4746

4847
### Additional subcommands
@@ -75,17 +74,15 @@ cp config.toml my-config.toml
7574
./target/release/baco scan --config my.toml --force # Force full rescan
7675
```
7776

78-
- **Phases**: 4 parallel (Indexing, Semgrep, CpgSlice, LlmStaticAnalysis) + sequential phases — some disabled by default (see [Configuration](docs/configuration.md))
77+
- **Phases**: 4 parallel (Indexing, Semgrep, CpgSlice, LlmStaticAnalysis) + 19 sequential phases — some disabled by default (see [Configuration](docs/configuration.md))
7978

80-
### What happens next
8179

82-
- **Phases**: 4 parallel (Indexing, Semgrep, CpgSlice, LlmStaticAnalysis) + sequential phases — some disabled by default (see [Configuration](docs/configuration.md))
8380

8481
## Features
8582

8683
- **Pipeline profiles**: `core` (default) runs essential phases; `all` enables experimental phases (still individually flag-gated) — set with `scanner.profile = "core" | "all"`
8784
- **Pipeline phases**: Indexing → Semgrep → CpgSlice (`[cpg]` section, requires Joern) → LlmStaticAnalysis → CweRouting (`router.enabled`) → RuleSynthesis (experimental) → LlmDiscovery → LlmVerification → Validate (opt-in) → SecurityAgentVerification (opt-in) → TicketCrossRef → GitAnalysis → CrossFileAnalysis → ConfidenceScoring → AiAggregation → ThreatModeling (`enable_threat_modeling`) → RootCauseDedup → MultiVerifier (`enable_multi_verifier`, experimental) → AutoPatching (`enable_auto_patching`, opt-in) → CveBootstrap → PocCompiler (`enable_poc_compilation`, opt-in) → ExploitSynth (`[exploit]` section, experimental) → VariantSearch → Reporting
88-
- **Parallel execution**: Indexing, Semgrep, CpgSlice, and LlmStaticAnalysis run concurrently; 20 sequential phases follow
85+
- **Parallel execution**: Indexing, Semgrep, CpgSlice, and LlmStaticAnalysis run concurrently; 19 sequential phases follow
8986
- **CWE-aware MoE (opt-in)**: BM25 RAG retrieval from CWE knowledge base, routes to specialized analysis paths — enable with `router.enabled = true`
9087
- **Research-backed**: 16 academic papers integrated (VulTriage, VulIn, MoCQ, MoEVD, AgentFlow) — see [Research Integration](docs/research-integration.md)
9188
- **Checkpoint/resume**: Crash recovery after each phase
@@ -134,12 +131,18 @@ cp config.toml my-config.toml
134131

135132
## Supported Languages
136133

137-
| Language | Static analysis | LLM analysis |
138-
| ---------- | ------------------------ | ------------ |
139-
| C / C++ | tree-sitter + semgrep | ✅ |
140-
| Rust | tree-sitter + semgrep | ✅ |
141-
| Python | tree-sitter + semgrep | ✅ |
142-
| JavaScript | tree-sitter + semgrep | ✅ |
134+
| Language | Static analysis | LLM analysis |
135+
| ------------- | ------------------------ | ------------ |
136+
| C / C++ | tree-sitter + semgrep | ✅ |
137+
| Rust | tree-sitter + semgrep | ✅ |
138+
| Python | tree-sitter + semgrep | ✅ |
139+
| JavaScript | tree-sitter + semgrep | ✅ |
140+
| TypeScript | tree-sitter + semgrep | ✅ |
141+
| PHP | tree-sitter + semgrep | ✅ |
142+
| Go | tree-sitter + semgrep | ✅ |
143+
| Java | tree-sitter + semgrep | ✅ |
144+
| C# | tree-sitter + semgrep | ✅ |
145+
| Ruby | tree-sitter + semgrep | ✅ |
143146

144147
## Outputs
145148

@@ -166,7 +169,7 @@ See [Research Integration](docs/research-integration.md) for per-paper details (
166169
- [Operator Tuning](docs/operator-tuning.md) — Performance flags and scenario-based tuning
167170
- [Output Interpretation](docs/output-interpretation.md) — Reading findings, confidence, triage verdicts
168171
- [Troubleshooting](docs/troubleshooting.md) — Common errors and fixes
169-
- [Roadmap](todo.md) — Completed and pending work
172+
170173

171174
### Reading Order
172175

@@ -179,7 +182,7 @@ Recommended for new users:
179182
6. **docs/operator-tuning.md** — performance tuning
180183
7. **docs/output-interpretation.md** — reading results
181184
8. **docs/troubleshooting.md** — error fixes
182-
9. **todo.md** — roadmap
185+
183186

184187
## Acknowledgements
185188

‎config.example.toml‎

Lines changed: 1 addition & 87 deletions
Original file line numberDiff line numberDiff line change
@@ -114,7 +114,7 @@ enable_hunt_prompts = false
114114
timeout_secs = 60
115115
# Maximum number of retries for failed requests (on 429/5xx)
116116
max_retries = 3
117-
# Retry backoff in milliseconds (exponential: ms * 2^attempt)
117+
# Retry backoff in milliseconds (linear: ms * (retries + 1))
118118
retry_backoff_ms = 2000
119119
# Maximum concurrent LLM requests (default: 4)
120120
max_concurrent = 4
@@ -190,58 +190,6 @@ project = "libxml2"
190190
# enabled = false # Set to true to enable the agent
191191
# max_turns = 10 # Maximum agent conversation turns per issue
192192

193-
# --- Paper-Integration Research Flags (P1-P5) ---
194-
# All flags default to disabled. Enable to opt in to experimental
195-
# research-backed analysis augmentations. See todo.md for details.
196-
197-
# VulTriage triple-path context augmentation — arXiv:2605.09461
198-
# Augments LLM input with control path (AST/CFG/DFG), knowledge path
199-
# (CWE pattern RAG), and semantic path (function summary) before judgement.
200-
#[vultriage]
201-
#enabled = false
202-
#control_path = true
203-
#knowledge_path = true
204-
#semantic_path = true
205-
206-
# VulnLLM-R policy-based generation — arXiv:2605.09461
207-
# Queries the LLM N times to build a CWE candidate set, then a final
208-
# call picks one label. Increases LLM cost ~5x.
209-
#[policy_sampling]
210-
#enabled = false
211-
#samples = 4
212-
213-
# VulnLLM-R agent scaffold — arXiv:2605.09461
214-
# Builds 3-path call-graph context + function-lookup tool per target.
215-
#[agent_scaffold]
216-
#enabled = false
217-
#max_rounds = 5
218-
#paths_per_target = 3
219-
220-
# MoCQ neuro-symbolic rule synthesis (RuleSynthesis 2.0) — arXiv:2605.13918
221-
# LLM proposes patterns in a DSL → symbolic validator gives feedback → loop.
222-
# Extends [scanner.rulesynth] with these fields:
223-
#[scanner.rulesynth]
224-
#mocq_mode = false
225-
#max_iterations = 5
226-
#corpus_path = "tests/fixtures/"
227-
228-
# PacVD primitive-API abstraction — arXiv:2605.07785
229-
# Appends callee abstraction at one of four granularity levels to the
230-
# LLM prompt. Level 1 = fuzzy branches only; 4 = concrete branches + key vars.
231-
#[pacvd]
232-
#enabled = false
233-
#level = 2
234-
#auto_level = false
235-
236-
# AgentFlow multi-agent harness synthesis — arXiv:2605.11835
237-
# Represents the harness as a typed graph DSL with a search loop.
238-
# Most invasive integration — static harness only until P5.5.
239-
#[agent_flow]
240-
#enabled = false
241-
#max_iterations = 10
242-
#requires_instrumented_target = false
243-
244-
245193
# CPG-guided slicing
246194
[cpg]
247195
# Whether CPG slicing is enabled
@@ -267,41 +215,7 @@ project = "libxml2"
267215
# Whether the Validate phase is enabled
268216
# enabled = false
269217

270-
# Triple-path context augmentation (P1: VulTriage)
271-
[vultriage]
272-
# Whether triple-path context augmentation is enabled
273-
# enabled = false
274-
# Whether to include the control path (AST/CFG/DFG verbalisation)
275-
# control_path = true
276-
# Whether to include the knowledge path (CWE pattern RAG)
277-
# knowledge_path = true
278-
# Whether to include the semantic path (function summary)
279-
# semantic_path = true
280-
281-
# Policy-based generation
282-
[policy_sampling]
283-
# Whether policy-based generation is enabled
284-
# enabled = false
285-
# Number of sampling rounds to build the policy
286-
# samples = 4
287218

288-
# Agent scaffold
289-
[agent_scaffold]
290-
# Whether the agent scaffold is enabled
291-
# enabled = false
292-
# Maximum interaction rounds per target function
293-
# max_rounds = 5
294-
# Number of call-graph paths to sample per target function
295-
# paths_per_target = 3
296-
297-
# Primitive-API abstraction
298-
[pacvd]
299-
# Whether PacVD abstraction is enabled
300-
# enabled = false
301-
# Abstraction level 1-4 (1=fuzzy, 4=concrete)
302-
# level = 2
303-
# Whether to auto-select the level based on the configured LLM model
304-
# auto_level = false
305219

306220
# Citation verification gate: deterministic checks that report citations
307221
# (file existence, line ranges) match the scanned tree before rendering.

‎docs/README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -10,7 +10,7 @@ Documentation for baco — a 23-phase LLM-assisted code scanner.
1010
|------|-------------|
1111
| [architecture.md](architecture.md) | The 23-phase pipeline, PhaseGraph, data flow |
1212
| [configuration.md](configuration.md) | All config options, LLM setup, phase flags |
13-
| [research-integration.md](research-integration.md) | The 20 papers integrated into baco |
13+
| [research-integration.md](research-integration.md) | The 31 papers integrated into baco |
1414
| [ci-integration.md](ci-integration.md) | CI/CD setup with SARIF output |
1515
| [llm-vuln-detection-papers-survey.md](llm-vuln-detection-papers-survey.md) | Full survey of 36 papers |
1616
| [example-report-screenshot.png](example-report-screenshot.png) | Sample HTML report preview |

‎docs/architecture.md‎

Lines changed: 7 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -11,9 +11,7 @@ BACO uses a **data-driven PhaseGraph** (`src/scanner/pipeline/orchestrator.rs`)
1111

1212
## Pipeline Phases
1313

14-
**Core Pipeline (23 phases):**
15-
16-
**Core Pipeline**:
14+
**Core Pipeline (23 phases total: 4 parallel + 19 sequential):**
1715

1816
> **Note:** Phase order is defined in `PhaseGraph::new()` (src/scanner/pipeline/orchestrator.rs:28-53). This table is manually maintained and should be updated when that code changes.
1917
@@ -36,12 +34,12 @@ BACO uses a **data-driven PhaseGraph** (`src/scanner/pipeline/orchestrator.rs`)
3634
| 15 | AI Aggregation | `llm.phases.aggregation` (API key present) |
3735
| 16 | Threat Modeling | `scanner.performance.enable_threat_modeling=false` |
3836
| 17 | Root Cause Deduplication | `scanner.performance.enable_root_cause_dedup=true` |
39-
| 19 | Auto-Patching | `scanner.performance.enable_auto_patching=false` |
40-
| 20 | CVE Bootstrap | `scanner.performance.enable_cve_bootstrap=true` |
41-
| 21 | PoC Compilation | `scanner.performance.enable_poc_compilation=false` |
42-
| 22 | Exploit Synthesis | `exploit.enabled=false` |
43-
| 23 | Variant Search | `scanner.performance.enable_variant_search=true` |
44-
| 24 | Reporting | Always-on |
37+
| 18 | Auto-Patching | `scanner.performance.enable_auto_patching=false` |
38+
| 19 | CVE Bootstrap | `scanner.performance.enable_cve_bootstrap=true` |
39+
| 20 | PoC Compilation | `scanner.performance.enable_poc_compilation=false` |
40+
| 21 | Exploit Synthesis | `exploit.enabled=false` |
41+
| 22 | Variant Search | `scanner.performance.enable_variant_search=true` |
42+
| 23 | Reporting | Always-on |
4543

4644
## Data Flow
4745

‎docs/llm-vuln-detection-papers-survey.md‎

Lines changed: 9 additions & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -26,11 +26,9 @@ This document is the **full survey** of 36 papers from the Awesome-LLMs-for-Vuln
2626
## 2. Top 5 Recommended Papers for Integration
2727

2828
All five papers were approved by the project owner on 2026-08-11 for
29-
integration into baco. The full implementation roadmap with sub-tasks,
30-
file paths, and acceptance criteria lives in [`../todo.md`](../todo.md)
31-
(P1-P5). Each integration is gated behind a config flag that defaults to
32-
`enabled = false`, so existing scanner behaviour is unchanged until an
33-
operator opts in.
29+
integration into baco (P1-P5). Each integration is gated behind a config
30+
flag that defaults to `enabled = false`, so existing scanner behaviour is
31+
unchanged until an operator opts in.
3432

3533
### P1 — VulTriage: Triple-Path Context Augmentation (arXiv:2605.09461)
3634

@@ -60,7 +58,7 @@ operator opts in.
6058
P1.5 Wire triple-path into prompt.
6159
- **Risk:** Low–medium. Reuses existing tree-sitter parsers and CPG slice
6260
from `CpgSlice` phase. Knowledge Path needs an embedding endpoint
63-
(open question Q1 in `todo.md`).
61+
(open question Q1).
6462
- **Why integrate:** Strong empirical false-positive reduction; minimal
6563
structural change to baco.
6664

@@ -102,7 +100,7 @@ operator opts in.
102100
retrieval tool, P2.5 Wire agent scaffold into SecurityAgentVerification.
103101
- **Risk:** Medium. Agent scaffold depends on call-graph quality and a
104102
tool-calling interface in `LlmClient` (shared prerequisite PS1 in
105-
`todo.md`). Policy sampling is 5× LLM calls — must be opt-in.
103+
Policy sampling is 5× LLM calls — must be opt-in.
106104
- **Why integrate:** Inference-time techniques are portable to any
107105
reasoning-capable model already configured in baco; agent scaffold
108106
gives the existing `SecurityAgentVerification` phase a concrete
@@ -134,7 +132,7 @@ operator opts in.
134132
corpus, P3.3 LLM proposer with feedback loop, P3.4 Emit accepted rules
135133
to disk, P3.5 Config + tests.
136134
- **Risk:** Medium. Needs a labelled trace corpus (open question Q2 in
137-
`todo.md`); start small (CWE-78, CWE-89).
135+
start small (CWE-78, CWE-89).
138136
- **Why integrate:** Bridges symbolic + LLM without replacing existing
139137
flows; reduces rule-author fatigue. Emitted Semgrep rules are consumed
140138
by the existing `Semgrep` phase with no further wiring.
@@ -183,7 +181,7 @@ operator opts in.
183181
walker, P4.3 Four-dimension extractor, P4.4 Level selector + prompt
184182
integration, P4.5 Model-aware level auto-selection.
185183
- **Risk:** Low–medium. Call-graph quality is the main dependency (shared
186-
prerequisite PS3 in `todo.md`).
184+
prerequisite PS3).
187185
- **Why integrate:** Largest reported precision boost among the five
188186
papers; abstraction is a strict superset of P1's Control Path, so the
189187
two compose.
@@ -237,9 +235,9 @@ operator opts in.
237235
P5.3 Runtime executor, P5.4 Diagnoser, P5.5 Proposer (search loop).
238236
- **Risk:** High. Coverage/sanitizer feedback requires build
239237
instrumentation that baco does not have today (shared prerequisite PS2
240-
in `todo.md`). Recommend shipping P5.1-P5.4 (static harness execution)
238+
Recommend shipping P5.1-P5.4 (static harness execution)
241239
first and deferring P5.5 (search loop) until instrumentation is
242-
available (open question Q3 in `todo.md`).
240+
available (open question Q3).
243241
- **Why integrate:** Most invasive but highest ceiling. Makes the
244242
existing `SecurityAgentVerification` phase a concrete, searchable
245243
harness space instead of a fixed pipeline; the typed DSL gives

‎docs/research-integration.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,10 @@
11
# Research-Backed Design
22

3-
Baco's LLM integration is informed by a survey of 36 papers from [Awesome-LLMs-for-Vulnerability-Detection](https://github.com/huhusmang/Awesome-LLMs-for-Vulnerability-Detection). This document details the 20 papers integrated into baco's architecture and what each contributes.
3+
Baco's LLM integration is informed by a survey of 36 papers from [Awesome-LLMs-for-Vulnerability-Detection](https://github.com/huhusmang/Awesome-LLMs-for-Vulnerability-Detection). This document details the 31 papers integrated into baco's architecture and what each contributes.
44

55
## Scope
66

7-
This document covers the **20 papers already integrated** into baco's architecture — what each contributes, where it's wired in, and the config flag that enables it.
7+
This document covers the **31 papers already integrated** into baco's architecture — what each contributes, where it's wired in, and the config flag that enables it.
88

99
## Integration Status Summary
1010

‎src/config/phases.rs‎

Lines changed: 1 addition & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -3,11 +3,7 @@ use std::path::PathBuf;
33

44
/// Aggregation configuration including false positive store settings
55
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
6-
pub struct AggregationConfig {
7-
/// Path to the false positive store JSON file
8-
#[serde(default)]
9-
pub fp_store_path: Option<PathBuf>,
10-
}
6+
pub struct AggregationConfig {}
117

128
/// Rule synthesis configuration (MoCQ: LLM→semgrep rule generation)
139
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -64,10 +60,6 @@ pub const DEFAULT_ENABLE_ROOT_CAUSE_DEDUP: bool = true;
6460
pub fn default_enable_root_cause_dedup() -> bool {
6561
DEFAULT_ENABLE_ROOT_CAUSE_DEDUP
6662
}
67-
pub const DEFAULT_ENABLE_MULTI_VERIFIER: bool = false;
68-
pub fn default_enable_multi_verifier() -> bool {
69-
DEFAULT_ENABLE_MULTI_VERIFIER
70-
}
7163
pub const DEFAULT_ENABLE_AUTO_PATCHING: bool = false;
7264
pub fn default_enable_auto_patching() -> bool {
7365
DEFAULT_ENABLE_AUTO_PATCHING

‎src/config/scanner.rs‎

Lines changed: 0 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -87,8 +87,6 @@ pub struct PerformanceSettings {
8787
pub enable_threat_modeling: bool,
8888
#[serde(default = "crate::config::default_enable_root_cause_dedup")]
8989
pub enable_root_cause_dedup: bool,
90-
#[serde(default = "crate::config::default_enable_multi_verifier")]
91-
pub enable_multi_verifier: bool,
9290
#[serde(default = "crate::config::default_enable_auto_patching")]
9391
pub enable_auto_patching: bool,
9492
#[serde(default = "crate::config::default_enable_poc_compilation")]
@@ -136,7 +134,6 @@ impl Default for PerformanceSettings {
136134
max_parallel_tasks: default_four(),
137135
enable_threat_modeling: crate::config::default_enable_threat_modeling(),
138136
enable_root_cause_dedup: crate::config::default_enable_root_cause_dedup(),
139-
enable_multi_verifier: crate::config::default_enable_multi_verifier(),
140137
enable_auto_patching: crate::config::default_enable_auto_patching(),
141138
enable_poc_compilation: crate::config::default_enable_poc_compilation(),
142139
enable_confidence_refinement: crate::config::default_enable_confidence_refinement(),

‎src/lib.rs‎

Lines changed: 0 additions & 21 deletions
Original file line numberDiff line numberDiff line change
@@ -1,24 +1,3 @@
1-
#![allow(
2-
clippy::items_after_test_module,
3-
clippy::assertions_on_constants,
4-
clippy::useless_vec,
5-
clippy::needless_return,
6-
clippy::match_single_binding,
7-
clippy::field_reassign_with_default,
8-
clippy::single_match,
9-
clippy::too_many_arguments,
10-
clippy::len_without_is_empty,
11-
clippy::redundant_clone,
12-
clippy::doc_markdown,
13-
clippy::borrow_deref_ref,
14-
clippy::len_zero,
15-
clippy::module_inception,
16-
clippy::vec_init_then_push,
17-
clippy::cloned_ref_to_slice_refs,
18-
clippy::needless_borrows_for_generic_args,
19-
clippy::test_attr_in_doctest
20-
)]
21-
221
pub mod agent;
232
pub mod agent_flow;
243
pub mod agent_scaffold;

0 commit comments

Comments
 (0)