A comprehensive collection of structured metadata from International Large-Scale Assessment (ILSA) research articles using machine learning methods. This repository contains extracted metadata from 130+ academic papers analyzing PISA, TIMSS, PIRLS, and other ILSA datasets.
The ilsa_survey_articles directory contains:
- JSON files: Structured metadata for 130+ research articles (extracted using AI/LLM pipeline)
- Database: SQLite database (
ilsa_knowledge_base.db) with all metadata - Parquet file: Tabular dataset (
ilsa_master.parquet) for analysis
Each JSON file follows the ILSAArticleMetadata schema with two main sections:
file_name: Source PDF filenametitle: Article titleauthors: List of authorsyear: Publication yeardoi: Digital Object Identifiervenue: Journal/conference namepublication_type: Journal, conference, etc.open_access: Accessibility statussource_category: Research type
survey_design: Weighting and sampling methodologysample_details: Sample size and country breakdownml_techniques: Machine learning algorithms usedconfounders_identified: Predictor variables (13 categories)main_findings: Structured results with performance metricsoutcome_summary: Narrative summary of findings
import pandas as pd
import json
from pathlib import Path
# Load all JSON files
json_dir = Path("ilsa_survey_articles/json")
articles = []
for json_file in json_dir.glob("*.json"):
with open(json_file) as f:
articles.append(json.load(f))
# Convert to DataFrame
df = pd.json_normalize(articles)import sqlite3
conn = sqlite3.connect("ilsa_survey_articles/ilsa_knowledge_base.db")
cursor = conn.cursor()
# List tables
cursor.execute("SELECT name FROM sqlite_master WHERE type='table';")
tables = cursor.fetchall()
print(tables)df = pd.read_parquet("ilsa_survey_articles/ilsa_master.parquet")The metadata was extracted using a custom pipeline:
- PDF Processing: PyMuPDF for text extraction
- LLM Extraction: OpenAI models for structured JSON extraction
- Schema Validation: Pydantic models for data quality
- Storage: SQLite and Parquet for analysis
Variables are categorized into 13 domains:
- Socioeconomic (ESCS, HOMEPOS, wealth)
- Demographic (gender, age, immigration)
- Student attitude (self-efficacy, motivation)
- Student behavior (study time, homework)
- Teacher (qualifications, experience)
- School (type, resources, climate)
- ICT (resources, computer use)
- Curriculum (type, instructional time)
- Parent/home (involvement, environment)
- Process data (aggregate task metrics)
- Prior achievement (test scores, grades)
- Peer effects (classroom climate)
- System level (GDP, education expenditure)
If you use this dataset, please cite the original research articles and acknowledge this collection.
The metadata extraction is provided for research purposes. Original article copyrights remain with their respective publishers.