A Flask web application that generates DCAT-AP CH compliant keywords by searching through multiple controlled vocabularies following the specified priority cascade.
-
Priority-based keyword search following DCAT-AP CH specifications:
- TERMDAT (Swiss federal terminology database)
- GEMET (Environmental terminology)
- Wikidata (General knowledge base)
- Literal keywords (fallback)
-
Modern web interface with responsive design
-
Real-time search with loading indicators
-
Structured results showing source, URI, and descriptions
-
Select and download keywords: Click on generated keyword tiles to select them and download the selected keywords in JSON format or upload them to an entry on I14Y.
-
Install Python dependencies:
pip install -r requirements.txt
-
Run the application:
python app.py
-
Open your browser and go to
http://localhost:5000
Compound keyword splitting relies on Hunspell lexicons. On DigitalOcean App Platform the deploy.sh script automatically installs hunspell plus the hunspell-de-CH, hunspell-fr, hunspell-it, and hunspell-en-gb packages so continuous deployment keeps working. For local development run:
sudo apt-get update && sudo apt-get install hunspell hunspell-de-ch hunspell-fr hunspell-it hunspell-en-gbYou can override the dictionary picked at runtime by setting HUNSPELL_DIC_PATH (absolute path to a .dic file) or HUNSPELL_LOCALE (e.g., de_CH). If no system dictionary is found the app falls back to the bundled vocabulary.
- Enter a search term in the input field.
- Click "Generate Keywords" or press Enter.
- Review the generated keywords sorted by priority.
- Click on keyword tiles to select them.
- Use the "Download Selected Keywords" button to download the selected keywords in JSON format.
- Use the provided URIs for DCAT-AP CH compliance
-
TERMDAT: Currently uses a placeholder implementation as TERMDAT doesn't have a public API. In a production environment, you would need to implement web scraping or use their specific API if available.
-
GEMET: Uses the GEMET API for environmental terminology searches. The current implementation includes a mock response structure that should be adapted based on the actual API response format.
-
Wikidata: Implements the Wikidata API for entity searches, returning proper Wikidata URIs.
-
Literal fallback: When no controlled vocabulary matches are found, the system falls back to literal keywords.
GET /- Main application interfacePOST /search- Keyword search endpoint- Request:
{"query": "search term"} - Response:
{"query": "...", "keywords": [...], "total": n}
- Request:
- Implement proper TERMDAT API integration
- Add caching for frequently searched terms (in progress)
- Include multi-language support
- Add export functionality for different formats
- Implement user preference settings for keyword sources
- Display search history in the web interface