SlopScan is a modern, production-grade technical auditor and codebase scanner designed to detect AI-generated "slop"βcode, commits, and documentation that look superficially correct but are structurally flawed, incomplete, redundant, or claim more than they actually deliver.
By combining robust static analysis (Abstract Syntax Trees) with LLM validation (Gemini 3.1 Flash-Lite & Qwen 72B), SlopScan identifies fake PR claims, generic AI commit patterns, hallucinated/outdated guides, and dead code before it pollutes your production codebase.
- Project Overview
- Core Features
- Tech Stack
- Architecture Overview
- Installation & Setup
- API Overview
- Project Structure
- Usage & Workflows
- Future Improvements
As generative AI coding assistants (like Copilots or coding agents) become deeply integrated into software workflows, codebases are experiencing an influx of high-volume, low-intent code contributions. This often leads to a new category of technical debt known as AI Slop:
- Misleading PR Descriptions: PR descriptions claiming to fix major bugs or add complex features when the underlying diff only contains superficial tweaks or string changes.
- Generic Commit Messages: Commit histories filled with vague AI-generated phrases (e.g.,
"Update index.js","Refactor code","Fix bug") without specific context or actual implementation matching the description. - Hallucinated Documentation: Readmes and guides containing obsolete version references, repetitive boilerplate, excessive marketing fluff, or references to functions/classes that do not exist in the code.
- Stubs & Placeholder Code: Production files left with empty function bodies (
pass), fake try-catch blocks returning static dummy successes ({ "status": "ok" }), hardcoded API secrets, or dead imports.
SlopScan empowers technical leads, open-source maintainers, and engineering judges by serving as an automated, zero-trust codebase auditor. It performs multi-layered static and semantic checks to guarantee that what developers (and their AI assistants) claim in descriptions matches the exact truth inside the repository files.
- Technical Reviewers & Judges: Speed up assessment of hackathon submissions or open-source PRs by verifying the truthfulness of claimed work.
- Maintainers & Tech Leads: Automate pre-merge audits to prevent low-quality code boilerplate or security vulnerabilities (secrets/localhost endpoints) from slipping in.
- Enterprise Teams: Ensure commit logs remain pristine, informative, and traceable.
Compares public GitHub pull requests against the actual underlying code changes.
- Diff Extraction: Fetches unified patches and diffs from GitHub's REST API.
- AI Claim Extraction: Gemini parses the PR description to extract testable claims (e.g., "Adds Auth middleware").
- Mismatch Detection: Evaluates each claim against a localized "change fingerprint" (which files were touched, if logic was modified, whether tests were included).
- Verdict Generation: Flags PRs as
TRUSTWORTHY,SUSPICIOUS, orMISLEADINGwith a structured confidence score.
Performs a deep review of a repository's recent commit history to catch generic and hallucinated commit messages.
- Diff Matching: Verifies the commit message against the exact code changes in that specific commit.
- Qwen Audit: Uses Qwen 2.5 72B Instruct to classify commits as
TRUSTWORTHY,GENERIC(vague or low detail), orHALLUCINATED(claiming features not in the diff). - Aggregate Quality Score: Computes a dynamic code-to-commit quality index for the repository.
Scans repository Markdown documents (.md and .mdx) to detect high-frequency AI patterns:
- Placeholder Finder: Flags unresolved template brackets,
coming soon,lorem ipsum, orTODO: write docs. - Marketing Fluff: Measures the density of generic buzzwords ("revolutionary", "seamless integration", "cutting-edge").
- Repetitive Paragraphs: Utilizes Jaccard similarity and NLTK/Levenshtein algorithms to locate semantic duplication across paragraphs.
- Claim Inconsistency: Qwen verifies all architectural and setup claims in the README against the actual source codebase.
Runs a deep file-by-file audit of codebases to locate structural issues, AI stubs, and anti-patterns:
- Regex Pattern Scorer: Identifies hardcoded keys, localhost development links, throw-not-implemented stubs, and empty stubs.
- Heuristic Scoring: Rates the severity of findings (low, medium, high, critical) and filters false positives in test or mock files.
- Deep AI Analysis: Concurrently calls Qwen 72B via OpenRouter to analyze long files or files with multiple regex hits, generating specific fixes.
- Core: Next.js App Router (React 19)
- Styling: Vanilla CSS + Tailwind CSS v4 custom theme extensions
- Icons:
lucide-react - Streaming Client: Native Web ReadableStreams consuming Server-Sent Events (SSE)
- Web Framework: FastAPI (Uvicorn / asyncio loops)
- Configuration: Pydantic Settings v2
- HTTP Clients:
httpx(async pools) - Database & Cache:
- Redis (metadata caching)
- Gemini API:
gemini-3.1-flash-lite(via google-generativeai SDK and direct betav1 HTTP endpoints) - OpenRouter:
qwen/qwen-2.5-72b-instruct(used for advanced codebase claims verification, commit verification, and deep file analysis)
SlopScan utilizes a modern direct-streaming, event-driven architecture centered around Server-Sent Events (SSE).
graph TD
A[Next.js Client] -->|HTTP POST Request| B[FastAPI Web Server]
B -->|Initiate Async Stream| C[Streaming Handler]
C -->|Fetch Source / Diffs| D[GitHub API]
C -->|Run Local Heuristics| E[Regex Scorer]
C -->|Request Structuring & Claims Verification| F[Gemini 3.1 Flash-Lite]
C -->|Request Deep Code Audit / Commits Review| G[OpenRouter Qwen 72B]
C -->|Stream SSE Chunk-by-Chunk| A
C -->|Cache Output| H[Redis]
- Frontend Initiation: The user pastes a GitHub URL into the Next.js input box.
- Stream Negotiation: The client establishes a
fetch()connection configured fortext/event-stream. - Async Processing Pipeline: The backend clones the repository (or fetches metadata/diffs via GitHub REST endpoints) and runs local heuristics.
- LLM Verification: Specific payloads are compiled and dispatched asynchronously using concurrent task managers.
- Real-time Progress Delivery: As phases complete, structured JSON progress packets (e.g.
{percent: 50, step: "Extracting claims"}) are streamed back over the HTTP pipe. - Result Presentation: The connection finishes with a final
resultpackage, rendering custom, interactive tabs on the dashboard.
- Python 3.10+
- Node.js 18+
- Redis Server (running locally or in the cloud)
- Git installed on the system path
- API Keys for Gemini and OpenRouter
-
Navigate to the server directory:
cd server -
Create and activate a virtual environment:
python -m venv venv # On Windows: venv\Scripts\activate # On macOS/Linux: source venv/bin/activate
-
Install python packages:
pip install -r requirements.txt
-
Configure environment variables: Create a
.envfile in the root of theserver/directory:# ββ Redis ββ REDIS_URL=redis://localhost:6379/0 # ββ GitHub Token ββ GITHUB_TOKEN=your_github_personal_access_token_here # ββ AI API Keys ββ GEMINI_API_KEY=your_gemini_api_key_here OPENROUTER_API_KEY=your_openrouter_api_key_here # ββ Settings ββ SANDBOX_ROOT=./tmp/sandboxed MAX_REPO_SIZE_MB=200 SCAN_TIMEOUT_SECONDS=120 # ββ CORS ββ CORS_ORIGINS=http://localhost:3000
-
Start the FastAPI server:
uvicorn main:app --host 127.0.0.1 --port 8000 --reload
The backend API will be available at
http://127.0.0.1:8000. You can inspect the interactive docs athttp://127.0.0.1:8000/docs.
-
Navigate to the client directory:
cd client -
Install Node packages:
npm install
-
Configure environment variables: Create a
.env.localfile in theclient/directory:NEXT_PUBLIC_API_URL=http://localhost:8000
-
Launch the development server:
npm run dev
Open
http://localhost:3000in your browser to view the application.
All major backend routes are mounted on the FastAPI application. The main routes are:
GET /github/repo?owner={owner}&name={name}: Fetches basic repository metrics (description, language, forks, stars) with Redis caching.GET /github/prs?owner={owner}&name={name}&state={state}: Lists the 20 most recent pull requests.GET /github/pr/{number}?owner={owner}&name={name}: Returns extensive details on a single PR, including code files changed, author metadata, and diff patches.GET /github/file?owner={owner}&name={name}&path={path}: Fetches raw file content from GitHub to enable in-app file viewing.
POST /api/pr-review/analyze: Evaluates a PR's description against its diff patches. Streams progress events and the final PR evaluation results.POST /api/commits/analyze: Scans recent commit messages against diffs using Qwen. Calculates aggregate quality and slop indexes.POST /api/docs/analyze: Clones the repository, parses all markdown files, checks paragraph Levinshtein similarity, and runs local stubs tests. Streams progress and outputs the executive summary with HTML Highlights.POST /api/code-review/analyze: Initiates a full codebase scan. Runs local regex patterns followed by OpenRouter deep audits on files that exceed length thresholds.
POST /api/code-review/summary: Sends a batch of code issues to Gemini to group, classify, and format into an executive report.GET /health: Verifies basic service connectivity.
server/
βββ core/
β βββ config.py # Configuration definitions & env loaders
β βββ redis.py # Redis client initialization & lifespan hooks
βββ routers/
β βββ github.py # Github fetching, details, and raw files
β βββ pr_review.py # Pull request auditing endpoint
β βββ commit_verify.py # Commit logging scanner
β βββ docs_verify.py # Markdown verifier endpoint
β βββ code_scan.py # Codebase scanner endpoint
βββ services/
β βββ github_service.py # GitHub REST API wrappers
β βββ pr_review_service.py# SSE pipeline for PR review
β βββ commit_verifier_service.py # SSE pipeline for commits
β βββ docs_verifier_service.py # SSE pipeline for docs
β βββ code_review_service.py # Codebase scanner coordinator
β βββ openrouter_service.py # Qwen completions integration
β βββ gemini_service.py # Gemini API endpoints integration
β βββ diff_service.py # Diff hunk parser & stats generator
βββ utils/
β βββ pattern_scorer.py # RegEx rule scoring engine for code quality
βββ main.py # FastAPI Application Entrypoint
client/
βββ src/
β βββ app/
β β βββ globals.css # Global stylesheets and Tailwind v4 theme variables
β β βββ layout.jsx # Google Sans font imports and metadata setup
β β βββ page.jsx # Landing page with repo search box
β β βββ repo/[owner]/[name]/
β β βββ page.jsx # Repo Dashboard card navigation
β β βββ prs/ # PR lists and detailing
β β βββ commits/ # Commit Verifier interface
β β βββ docs/ # Docs Verifier dashboard
β β βββ scan/ # Code Scanner file explorer tree
β βββ components/
β β βββ repo/ # RepoDashboard and Sticky Navigation
β β βββ pr/ # PRLists, Commit details, and Verdict panels
β β βββ commits/ # CommitVerifier component UI
β β βββ docs/ # DocsVerifier panel UI
β β βββ scanner/ # IDE-like CodeViewer, FileExplorer tree, and Issue panels
β β βββ ui/ # Badges, Cards, Spinners, and Progress Streams
β βββ hooks/
β β βββ useActionStream.js # Core hook consuming SSE text streams
β βββ lib/
β βββ api.js # Fetch wrappers for API endpoints
β βββ constants.js # Global dictionary for status codes and colors
β βββ github.js # GitHub repository URL parsing utility
This project is licensed under the MIT License - see the LICENSE file for details.