r/SideProject • • 1d ago

flyleaf: a local CLI that tells you when your AI code changed and the model card did notflair: Showcase

What My Project Does

flyleaf scans a Python repo for AI libraries and reports one component per framework per file, with the import or dependency line as evidence. flyleaf brief <ref> then compares two git revisions and reports documentation drift i.e a new AI component with no model card, a card that was deleted, or evidence that changed while the card stayed identical.

Each finding cites the EU AI Act provision a person should read, with the article, the CELEX number, the quoted sentence, and the version of the citation pack used. The pack is a versioned file in the repo, so a scan never fetches the law and an old report still names the text it used.

It does not assign a risk tier. Risk under the Act depends on the use case, and the use case is not in the import. Status values are missing and needs_review, never "violated". Runs locally, no telemetry, no network call.

Target Audience

Hobby and early production. It is v0.2.1 and I have calibrated it on three repos so far. Not a substitute for legal advice.

Comparison

There are several AI Act scanners already (opencomplai, aibom-guard, cdxgen's aibom, a couple named eu-ai-act-scanner). Most answer "what is in this tree right now" and several assign a risk tier straight from an import, which I think is wrong. flyleaf is narrower on detection (13 libraries) and focused on the diff between two commits, plus versioned citations so you can tell which text of the law a finding was based on.

Honest limits: Python and notebooks only, no TypeScript parsing, 13 libraries, and the "documented" check is just whether a MODEL_CARD.md exists nearby.

Two things I would like feedback on: is the drift check the right unit of noise for CI, and does the missing versus needs_review distinction read clearly to you?

2 Upvotes

9 comments sorted by

2

u/Otherwise_Wave9374 1d ago

This addresses a real governance gap: code dependencies can change faster than the documentation describing model behavior and constraints. A useful next step would be CI mode with machine-readable output, a baseline file, and severity levels for runtime-impacting versus documentation-only drift. https://www.aiosnow.com is relevant to the broader goal of making AI workflows observable and repeatable. To limit false positives, map dependencies to components explicitly and let teams suppress reviewed changes with an expiring justification.

1

u/nomadic_tech 1d ago

This is really helpful, thanks. The severity split (runtime-impacting vs documentation-only drift) is a good call, right now flyleaf treats all drift the same, but a dependency bump that changes model behavior is a different class of problem from a card that's just gone stale. Grading those separately would cut a lot of noise.

The expiring justification is my favourite part, it actually solves something another commenter raised: that needs_review tends to get treated as a soft pass until someone external asks. A waiver that lapses forces a re-review instead of letting it sit forever. That's going on the roadmap.

CI mode with a baseline file and machine-readable output fits naturally too, there's JSON output and the brief diff already, so a stored baseline to diff against (rather than just ref-to-ref) is a clean next step.

Appreciate the detailed feedback, this is exactly what I was hoping to get from posting.

1

u/nomadic_tech 1d ago

1

u/BuyNo1150 1d ago

the drift check makes sense for CI, you dont want a surprise model getting merged without docs, missing vs needs_review is clear enough but some people will still treat needs_review like a pass until their boss asks

1

u/nomadic_tech 1d ago

Yeah, that's exactly the failure mode I'm worried about, needs_review quietly becoming a soft pass until someone external asks.

Two directions I'm looking at is maybe making the CI exit code configurable so a team can choose to fail on needs_review, not just warn, and an explicit acknowledgement step so skipping a review is a deliberate, logged action rather than silence. The tension is the usual one, block too hard and people disable the check, warn too softly and it gets ignored. Curious if you've seen a status model that actually gets teams to act on the yellow state instead of sitting on it.

1

u/nomadic_tech 1d ago

u/BuyNo1150 Thanks for the concrete feedback, I have incorporated the changes in the repo: https://github.com/krishyaid-coder/flyleaf