Qwen2.5-Coder-7B — Solidity Audit (V3, merged, deploy-ready)

Full-weight merge of the V3 QLoRA adapter into Qwen/Qwen2.5-Coder-7B-Instruct (bf16, 7.6B params, single 15.2 GB safetensors shard). This repo exists specifically to be deployed on a Hugging Face Inference Endpoint (or served with TGI/vLLM) — the adapter version is for PEFT users.

Part of the Solidity Audit Scanner project.

What it does

Given a Solidity snippet, produces a structured audit report: the primary finding classified into exactly one of 8 canonical categories (Reentrancy, Arithmetic Error, Access Control Flaw, Input Validation Error, Frontrunning / MEV, Denial of Service (DoS), Business Logic Violation, Clean (No Vulnerability)) — with functionality analysis, description, and remediation recommendation.

Measured evaluation (n=100 stratified held-out, greedy, seed 42)

Metric V3 (this model) V2 V1
Detection precision 0.986 0.986 1.00
Detection recall 0.986 0.971 1.00
False-positive rate on clean 0.033 0.033 0.00
Finding-type match (exact canonical) 0.449 0.485 0.30 (lenient)

Full per-row records: V3 eval_results.json.

Honest read: V1's perfect detection was a dataset artifact (see its card); V3 is the best detection config with the cleanest labels; the type-match gap between V2/V3 is within sampling noise at n≈70, and residual type errors are semantic ambiguity in reference labels, not model bugs.

Usage (OpenAI-compatible once deployed on an endpoint; transformers locally)

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "Offlin33er/qwen25-coder-7b-solidity-audit-v3-merged", dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained("Offlin33er/qwen25-coder-7b-solidity-audit-v3-merged")

messages = [
    {"role": "system", "content": SYSTEM_PROMPT_V3},  # the V3 taxonomy prompt from the dataset card
    {"role": "user", "content": "Audit the following Solidity code for security vulnerabilities.\n\n```solidity\n<YOUR CODE>\n```"},
]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

The exact SYSTEM_PROMPT_V3 text is on the dataset card — it was trained verbatim and must be used as-is.

Deployment

  • HF Inference Endpoint: 1× Nvidia L4 (24 GB) recommended (~$0.80/hr, scale-to-zero for near-zero idle cost). A T4 (16 GB) does not fit bf16.
  • Greedy decoding, 512 max new tokens matches the evaluated configuration.

Training

1 epoch (139 steps) over Offlin33er/solidity-audit-instruct-v3 (2,224 rows, canonical 8-class taxonomy, split-preserving vs V1/V2); QLoRA r=32 on all attn+MLP projections, LR 2e-4 cosine, merged after training. Tracked on the trackio dashboard.

Limitations

  • Trained on ~2.5K Solodit-audited snippets; ~290 one-off label types in the source taxonomy.
  • ~Half of detected findings still get an adjacent canonical label (44.9% exact match) — the residual errors are reference-label ambiguity (documented in the V3 card).
  • Not a substitute for a professional audit. Findings are candidates for human review.

Defensive security tooling: analyzes code you paste in. Only audit contracts you are authorized to review.

Downloads last month
319
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Offlin33er/qwen25-coder-7b-solidity-audit-v3-merged

Base model

Qwen/Qwen2.5-7B
Finetuned
(475)
this model

Dataset used to train Offlin33er/qwen25-coder-7b-solidity-audit-v3-merged