How to set up, code, test, review, and release so contributions meet our Definition of Done.
All contributors must follow the Oregon State University Student Code of Conduct and the team’s charter agreement.
- Treat all collaborators with respect and professionalism.
- Provide decent participation during meetings and reviews.
- Raise the issue privately with the team first.
- Issues of academic or ethical concern should be reported directly to the instructor.
- Report any inappropriate or unprofessional behavior to the TA, instructor or project manager.
Owner: Bradley Rule Next Review: 05/15/26
Pipeline diagram:
documentation/architecture.png- visual overview of the full extraction pipeline.
-
- Python 3.10 +
- pip installed
- Access to GitHub repository
- Ollama installed and running locally
- Minimum hardware: 8 GB RAM (16 GB recommended for
qwen2.5:7b) - Pull the required models before running the classify/extract pipeline:
ollama pull qwen2.5:7b # default extraction model (~5 GB) - Verify Ollama is running:
ollama list
- Minimum hardware: 8 GB RAM (16 GB recommended for
- Tesseract OCR — required for scanned PDFs
- Windows: download and run the installer from UB Mannheim, or
choco install tesseract - macOS:
brew install tesseract - Ubuntu/Debian:
sudo apt install tesseract-ocr - After install, ensure
tesseractis on your PATH:tesseract --version
- Windows: download and run the installer from UB Mannheim, or
- Ghostscript — optional; improves table extraction on bordered PDFs
- The pipeline works without it. Ghostscript is only used by camelot's lattice mode, a last-resort fallback for bordered tables that PyMuPDF and camelot stream mode both missed.
- Windows: download from ghostscript.com, or
choco install ghostscript - macOS:
brew install ghostscript - Ubuntu/Debian:
sudo apt install ghostscript
git clone https://github.com/NovakLabOSU/FracFeedExtractor.git
cd FracFeedExtractor
python -m venv venv
source venv/bin/activate
# Windows: venv\Scripts\activate
pip install -e ".[dev]"
-
- If you do have the dataset downloaded locally on your machine:
python scripts/full_pipeline.py --local <file_path_to_dataset>- Note: Your local dataset folder should be formatted like so:
folder |-> useful |-> pdfs |-> ... |-> not-useful |-> pdfs |-> ...
- Note: Your local dataset folder should be formatted like so:
- If you do not have the dataset downloaded or cannot have it on your locally on your machine:
python scripts/full_pipeline.py --api- Note: You will need access to the .env file
- If you do have the dataset downloaded locally on your machine:
-
Use
src/pipeline/classify_extract.pyto classify PDFs and extract structured diet data in a single step. Requires trained model artifacts insrc/classifier/models/(run the full pipeline first, or see Retraining the Classifier below).# Single PDF python src/pipeline/classify_extract.py path/to/file.pdf # Folder of PDFs (sequential) python src/pipeline/classify_extract.py path/to/pdfs/ # All options python src/pipeline/classify_extract.py path/to/pdfs/ \ --model-dir src/classifier/models \ --llm-model qwen2.5:7b \ --output-dir data/results \ --confidence-threshold 0.70 \ --max-chars 12000 \ --num-ctx 4096 \ --workers 4
Flag Default Description --model-dirsrc/classifier/modelsDirectory containing classifier artifacts --llm-modelqwen2.5:7bOllama model for extraction --output-dirdata/resultsDestination for JSON results and summary CSV --confidence-threshold0.70Probability threshold for "useful" classification --max-chars12000Maximum characters sent to the LLM --num-ctx4096Ollama context window size (tokens) --workers1Parallel worker processes ( 1= sequential) -
Each PDF classified as "useful" produces a JSON file in
data/results/metrics/:{ "source_file": "Smith_2002.pdf", "extracted_at": "2026-04-24T14:32:00", "metrics": { "species_name": "Esox lucius", "study_location": "Lake Windermere, UK", "study_date": "1998-2000", "num_empty_stomachs": 42, "num_nonempty_stomachs": 158, "sample_size": 200, "fraction_feeding": 0.79 } } -
- Sensitive information such as API keys will be stored in a local .env file which will be excluded by .gitignore.
- Never hardcode secrets
Owner: Raymond Cen Next Review: 05/15/26
We will use the feature-branch workflow with all merges handled through PRs.
- Default branch:
main - Branch naming template:
feature/short-namebugfix/short-name
- Rebase your working branch with main, and often, before submitting a PR (simpler conflict resolution)
Owner: Zahra Zahir Ahmed Alsulaimawi Next Review: 05/15/26
Issue titles should start with the following tags to designate intent:
FEAT:New feature request.- Include problem and feature description within issue description field.
- Include note of requirement ID where applicable (ex:
Requirement ID: REQ-005).
BUG:Bug report.- Include problem description.
DOC:Documentation changes.- Brief description of what needs to be changed or added.
Example:
FEAT: Example Issue
Requirement ID: REQ-XXX
Problem Desciption:
- Brief explanation of why this is important
Feature Description:
- Brief explanation of expected implementation
Owner: Zahra Zahir Ahmed Alsulaimawi Next Review: 05/15/26
We will use the Conventional Commit format for clarity and traceability
<type>(scope): short summary [issue number if applicable]
[optional body]
[optional footer]
Examples:
feat(parser): implement CSV input parsing
fix(ci): update pytest command in workflow [#42]
docs(readme): add setup section
Owner: Raymond Cen Next Review: 05/15/26
We use Black for automatic code formatting and Flake8 for linting to maintain consistent style and prevent common Python errors.
-
- Config file:
pyproject.toml - Install
pip install black - Local usage:
# Check formatting without changing files black --check src tests # Automatically reformat code black src tests
- Config file:
-
- Config file:
pyproject.toml - Install
pip install flake8 - Local usage:
flake8 src/ tests/
- Configured to ignore line length violations (E501) and other minor style differences.
- Config file:
Owner: Sean Clayton Next Review: 05/15/26
-
pytest
-
# Run all tests pytest tests/ # Run tests with coverage coverage run -m pytest tests/ coverage report -m coverage html
Owner: Sean Clayton Next Review: 05/15/26
-
- New features must include unit or integration tests.
- Coverage thresholds: aim for >80% for core modules; all critical paths must be tested.
- Tests must pass locally before creating a PR.
-
- Keep PRs focused and small (<300 lines changed if possible).
- Include related issue references in the PR description.
- Clearly describe what the PR changes and why.
-
- At least one approving review is required for non-trivial changes.
- Reviewers check code quality, tests, CI status, and adherence to style guides.
- Ensure all linting and formatting checks pass.
-
- CI must pass all mandatory jobs before merging.
- PRs should be rebased on the latest
mainbranch before merge if there are conflicts.
Owner: Bradley Rule Next Review: 05/15/26
Continuous integration ensures all contributions meet quality standards automatically.
-
- GitHub Actions workflow:
.github/workflows/pdf_extraction_ci.yml
- GitHub Actions workflow:
-
- Install dependencies
- Code style checks: Black & Flake8
- Unit tests & coverage
- Functional validation: PDF extraction tests
-
- Navigate to the Actions tab in GitHub.
- Select the workflow run and expand individual jobs to see logs.
-
- All jobs must complete successfully.
- Any failing test, linter, or formatting check blocks the merge.
- Artifacts (e.g., coverage reports) are uploaded automatically and can be reviewed.
Owner: Sean Clayton Next Review: 05/15/26
State how to report vulnerabilities, prohibited patterns (hard-coded secrets), dependency update policy, and scanning tools.
- Never commit sensitive credentials, tokens, or API keys.
- Secrets are stored locally in .env files and excluded via .gitignore.
- Dependencies are managed in
pyproject.toml. - Use pip-audit monthly to check for vulnerabilities.
- Security issues or potential breaches should be reported privately to the Project Manager and TA.
Owner: Raymond Cen Next Review: 05/15/26
README.md: Provides an overview of the project and its goals. It should also detail where to find information and contact individuals.documentation/: This directory will store all documentation for the program and its features including standards and usage information.Comments:- General comments should be kept short and to the point.
- Functions should contain a docstring as the first statement within a function providing a brief explanation of the function and a quick explanation of parameters and return values.
- Inline comments should be reserved for places where the function of code is difficult to understand or infer.
Owner: Zahra Zahir Ahmed Alsulaimawi Next Review: 05/15/26
- We follow semantic versioning**:
MAJOR.MINOR.PATCH- MAJOR: breaking changes
- MINOR: new features, backwards-compatible
- PATCH: bug fixes, minor improvements
- Example:
1.2.0→1.2.1(patch),1.3.0(minor),2.0.0(major)
- Each release must be tagged in Git using the version number:
git tag -a v1.2.0 -m "Release v1.2.0: <short description>"
git push origin v1.2.0-
For each release, include:
- Added: new features
- Changed: updates or improvements
- Fixed: bug fixes
Example entry:
## [1.2.0] - 2025-11-10
### Added
- New feed extraction method
### Fixed
- Timeout handling in API fetch
- Before publishing, ensure all CI/CD pipelines pass.
- Prepare the release branch or merge main into it.
- Include updated documentation and changelog.
- In case of a faulty release
- Revert the release tag in Git:
git tag -d v1.2.0 git push origin :refs/tags/v1.2.0- Revert the merge commit on the main branch:
git revert -m 1 <merge-commit-hash> git push origin main- Update CHANGELOG to reflect the rollback.
- Notify the team and project partner of the rollback.
Owner: Bradley Rule Next Review: 05/15/26
The classifier artifacts are saved in src/classifier/models/. To retrain with new or updated labeled data:
-
Add labeled text files to
data/processed-text/and updatedata/labels.jsonwith"filename.txt": "useful"or"filename.txt": "not useful"entries. -
Run the trainer directly:
python -m src.classifier.train_model
This reads from
data/processed-text/anddata/labels.json, trains a TF-IDF + XGBoost model, and saves three artifacts:src/classifier/models/pdf_classifier.json- XGBoost modelsrc/classifier/models/tfidf_vectorizer.pkl- TF-IDF vectorizersrc/classifier/models/label_encoder.pkl- LabelEncoder
-
Or run the full pipeline, which trains the model as a final step:
python scripts/full_pipeline.py --local <path_to_dataset>
Key tunable parameters in src/classifier/train_model.py:
max_featuresinTfidfVectorizer(default: 10,000)eta,max_depth,subsamplein the XGBoostparamsdictearly_stopping_rounds(default: 20)
Extraction fields are defined in two places:
-
src/extraction/models.py- thePredatorDietMetricsPydantic model. Add a new optional field with the appropriate type and aNonedefault:prey_taxa: Optional[list[str]] = None
-
src/extraction/llm_client.py- the system prompt that instructs the LLM. Add a description of the new field and its expected format to the prompt string. -
src/pipeline/classify_extract.pyandsrc/pipeline/extract_from_txt.py- update therowdict andfieldnameslist in the summary CSV writer to include the new column.
After adding a field, run pytest tests/test_llm_text.py to verify that the prompt changes do not break existing extraction tests.
- Primary Communications: Slack and Teams
- Meetings: Fridays 1 PM PST
- Project Partner: Mark Novak, Fridays 8:30AM PST (biweekly check-ins)
- TA Meetings: Thursdays 1:30PM PST
Owner: Zahra Zahir Ahmed Alsulaimawi Next Review: 05/15/26