Code-Projekte & Repositories
Aktive Entwicklungsprojekte via GitHub. Mein offenes Backoffice für kollaborative Wissenschaft.
EmotiView
Sprache: Jupyter Notebook ⭐ 1
Zuletzt aktualisiert: 2026-08-18
README anzeigen
EmotiView: Neural-Autonomic Synchrony and Embodied Integration
This repository accompanies ongoing research investigating the dynamic interplay between neural activity and autonomic nervous system responses during emotional experiences. Here you'll find the research article, presentations, analysis pipeline, and results—updated in real-time as the project progresses.
Principal Investigator: Cagatay Özcan Jagiello Gutt
| Platform | Role | Contents |
|---|---|---|
| OSF | Research output | Article, documentation |
| GitHub | Technical implementation | Analysis pipeline, results, presentations, proposal |
Research Abstract
Emotional states are fundamentally embodied, emerging from the dynamic interplay between central neural processing and peripheral physiological adjustments orchestrated by the autonomic nervous system (ANS). While ANS outputs like heart rate variability (HRV) and electrodermal activity (EDA) reflect emotional arousal and valence, understanding the precise temporal coordination between brain activity and these peripheral signals is crucial for elucidating brain-body interactions. This study investigates neural-autonomic phase synchrony during the conscious processing of distinct emotional states (positive, negative, neutral) by quantifying the temporal alignment between cortical and physiological rhythms.
We employ a multimodal approach, simultaneously recording high-temporal-resolution electroencephalography (EEG), electrocardiography (ECG) for HRV analysis (specifically Root Mean Square of Successive Differences, RMSSD), EDA, and functional near-infrared spectroscopy (fNIRS) while participants view validated emotional video clips. Our primary analysis quantifies the Phase Locking Value (PLV) between frontal EEG oscillations (Alpha, Beta bands) and continuous signals derived from HRV (reflecting parasympathetic influence) and phasic EDA (reflecting sympathetic influence). EEG channel selection for PLV analysis is informed by task-related hemodynamic activity measured via fNIRS to focus on functionally relevant cortical areas.
We hypothesize that PLV, indicating brain-body temporal integration, will be significantly modulated by emotional content compared to neutral conditions. We further expect synchrony strength to correlate with subjective arousal ratings. By examining the phase synchrony between brain signals and ANS-mediated physiological outputs, this research provides novel insights into the dynamic, embodied mechanisms underlying emotional experience. Understanding this temporal binding is critical for models of psychophysiological function and may inform assessments of cognitive load or stress regulation capacity.
Core Research Aims & Hypotheses
This project seeks to understand how the brain and body coordinate during emotional processing, focusing on neural-autonomic phase synchrony. Key hypotheses include:
- Emotional Modulation of Synchrony: Neural-autonomic synchrony (Phase Locking Value - PLV) will be enhanced during the processing of positive and negative emotional stimuli compared to neutral stimuli, for both brain-heart (EEG-HRV) and brain-sudomotor (EEG-EDA) coupling.
- Synchrony and Subjective Arousal: The magnitude of neural-autonomic synchrony will positively correlate with subjective ratings of emotional arousal during emotional conditions.
- Baseline Vagal Tone and Task-Related Synchrony: Individual differences in baseline parasympathetic regulation (resting-state RMSSD) will be associated with the degree of EEG-HRV synchrony during negative emotional stimuli.
- Frontal Asymmetry and Branch-Specific Synchrony: The direction of prefrontal cortical asymmetry (Frontal Asymmetry Index - FAI) will be differentially associated with the strength of phase synchrony involving distinct autonomic branches (EEG-HRV vs. EEG-EDA).
For a comprehensive understanding of the theoretical background, detailed methodology, and specific work packages, please refer to the full proposal document.
Methodology Overview
A multimodal experimental design is employed, involving:
- Stimuli: Standardized, emotionally evocative video clips (positive, negative, neutral) from the E-MOVIE database.
- Participants: Healthy young adults, screened for relevant criteria.
- Data Acquisition: Simultaneous recording of:
- Electroencephalography (EEG): To measure prefrontal neural dynamics.
- Functional Near-Infrared Spectroscopy (fNIRS): To localize hemodynamic activity in prefrontal and parietal regions, informing EEG channel selection.
- Electrocardiography (ECG): For Heart Rate Variability (HRV) analysis.
- Electrodermal Activity (EDA): To measure sympathetic nervous system activity.
- Subjective Measures: Self-Assessment Manikin (SAM) for valence and arousal, Positive and Negative Affect Schedule (PANAS), and Behavioural Inhibition/Approach System (BIS/BAS) scales.
Repository Contents
Research Output (OSF)
- Article: (Coming soon) The research article summarizing findings and contributions.
Technical Implementation (GitHub)
EV_results/: Processed data, analysis metrics, and visualizations.EV_analysis/: The Nextflow-based analysis pipeline with Python modules.EV_presentation/: Slides and presentation materials.EV_proposal/: The original research proposal with methodology and analysis plan.
Analysis Pipeline
The analysis pipeline is built on the AnalysisToolbox—a modular Nextflow framework for scalable, reproducible data processing with automatic result synchronization. The EmotiView-specific pipeline in EV_analysis/ extends this framework to:
- Load and parse multi-modal raw data (EEG, fNIRS, ECG, EDA, questionnaires).
- Perform standardized preprocessing steps specific to each physiological modality.
- Extract key features and metrics (e.g., EEG power, FAI, RMSSD, fNIRS ROI activation, PLV).
- Generate participant-level results and aggregated summaries.
Configuration is managed via EV_analysis/EV_parameters.config. See the AnalysisToolbox documentation for framework details.
Project Status
Data collection and Thesis writing.
Contributors
| Name | Role | Contact |
|---|---|---|
| Cagatay Özcan Jagiello Gutt | Principal Investigator | |
| Ben Gopin | Technical Assistant | |
| Gerrit Jostler | Technical Assistant |
AnalysisToolbox
Sprache: Python
Zuletzt aktualisiert: 2026-09-02
README anzeigen
AnalysisToolbox
A modular framework for automated data processing and statistical analysis pipelines. Built on Nextflow for scalable, reproducible workflows with automatic result synchronization.
Overview
The AnalysisToolbox provides infrastructure for building data processing pipelines that:
- Process multiple datasets in parallel with automatic participant discovery
- Handle diverse data types through a generic reader/processor/analyzer architecture
- Track progress via per-participant logging visible live in the web UI
- Recover gracefully from failures without losing completed work
The framework is domain-agnostic — modules follow simple input/output conventions (Parquet/FIF files) and can implement any processing logic.
Repository Structure
AnalysisToolbox/
├── gitatbx/ # pip-installable package (installed via pip install GitAtbx)
│ ├── bin/ # workflow_wrapper.nf, log_to_parquet.py, nextflow.config, ...
│ ├── modules/ # analyzers/, processors/, readers/, utils/
│ ├── templates/ # workflow_template.nf, modules_template.nf, parameters_template.config
│ └── utils/ # serve_html.ps1, reinject.sh, result_collector.py
├── pyproject.toml
└── README.md
On first run gitatbx creates a symlink ~/Documents/GitAtbxModules → <site-packages>/gitatbx/modules/. The symlink always reflects the live installed version — upgrading via pip install --upgrade GitAtbx automatically shows updated modules through the same symlink path.
Prerequisites
Java Runtime (required by Nextflow)
sudo apt update && sudo apt install default-jre
java -versionNextflow
curl -s https://get.nextflow.io | bash
sudo mv nextflow /usr/local/bin/
nextflow -versionInstall
pip install GitAtbx
Python dependencies (numpy, scipy, polars, mne, neurokit2, …) are installed automatically.
On first run gitatbx creates a symlink ~/Documents/GitAtbxModules → <site-packages>/gitatbx/modules/ and saves the path to ~/.gitatbx_config.
Usage
Commands
| Command | What it does |
|---|---|
gitatbx init <dir> | Scaffold a new analysis project |
gitatbx run [pattern] | Find and run a pipeline by project name pattern |
gitatbx serve [--dir DIR] [--port PORT] | Serve results HTML locally in browser |
gitatbx reinject <PID> [options] | Reinject a corrected output for one participant |
gitatbx move <dest> | Move the deployed modules folder to a new location |
gitatbx config show | Print current configuration (~/.gitatbx_config) |
gitatbx init
Prompts for project name, raw data directory, Python executable, toolbox path, git author identity, and an optional GitHub remote URL for the results repo. Then creates:
<dir>/
├── {name}_analysis/
│ ├── {name}_pipeline.nf ← edit your workflow here
│ ├── {name}_modules.nf ← add IOInterface includes here
│ └── {name}_parameters.config ← pre-filled paths, params, and git identity
└── {name}_results/
└── .git/ ← initialised + remote added (if URL provided)
Git author name and email default to your global git config values if already set. The remote URL is validated immediately with git ls-remote — if authentication fails (e.g. SSH key not yet added to GitHub), a warning is printed with a link to the GitHub SSH setup guide.
The pipeline uses the stamped params.git_user_name / params.git_user_email as the commit author for all automatic result syncs.
gitatbx run
GitAtbx searches the entire accessible filesystem (home directory and all drives on Windows) for a directory named (name)_analysis containing a *_pipeline.nf, then runs it automatically. No need to cd anywhere.
gitatbx run (name) --resume # continue a previous run
gitatbx run # no pattern: use current directory
Found paths are cached in ~/.gitatbx_config so subsequent calls are instant.
gitatbx serve
Starts a local HTTP server to browse results HTML generated by the pipeline.
gitatbx serve --dir ../EV_results --port 8080gitatbx reinject
Places a corrected parquet into corrections/<script_name>/, marks the participant for replay, invalidates relevant Nextflow cache entries, and resumes the pipeline for that participant only.
gitatbx reinject EV_002 --corrected-file fixed.parquet --script-name filtering_processorgitatbx move
To move the symlink to a custom location:
gitatbx move /mnt/d/repoShaggy/GitAtbxModules
gitatbx moves the folder and updates ~/.gitatbx_config automatically. gitatbx init will then default toolbox_dir to that path when scaffolding new projects.
Key Components
bin/workflow_wrapper.nf
Discovers participant directories, manages per-participant output folders, runs log_to_parquet.py and interactive_plotter.py automatically, and handles per-participant git sync on completion.
IOInterface
Generic Nextflow process that runs any script (reader/processor/analyzer) with automatic logging. Every module in modules/ is called through IOInterface.
Modules
The modules/ folder is a curated but non-exhaustive starting collection of commonly useful scripts. The first time gitatbx is run it mirrors the full collection to ~/Documents/GitAtbxModules (Windows) or ~/GitAtbxModules (Linux/macOS) automatically. You can move that folder anywhere with gitatbx move. The local copy is yours to extend: add domain-specific scripts, modify existing ones, or organise them into subfolders. Any script there can be called via IOInterface identically to the built-in modules, as long as it follows the same convention: positional CLI arguments in, Parquet or FIF outputs, non-zero exit on failure.
Authors
Cagatay Özcan Jagiello Gutt — Lead Developer ORCID: https://orcid.org/0000-0002-1774-532X
5ha99y
Sprache: JavaScript
Zuletzt aktualisiert: 2026-07-02
README anzeigen
Zola GitHub Pages Site - Scientific Hub
A static website built with Zola that automatically syncs content from your scientific profiles.
Website URL
https://cgutt-hub.github.io/5ha99y
How It Works
Automatic Updates
The site automatically pulls data from:
- GitHub — Your repositories and code projects
- ORCID — Publications and works
- OSF — Research projects and data (if configured)
Deployment Flow
1. Push to main branch
↓
2. GitHub Actions runs
↓
3. Fetches data from APIs
↓
4. Builds site with Zola
↓
5. Deploys to gh-pages branch
↓
6. Website updates automatically!Repository Structure
Source Files (Edit These)
config.toml- Site configurationcontent/- Your content (CV, contact, welcome post, etc.)templates/- HTML templatesstatic/- Static assets (CSS, images)scripts/fetch_data.py- Fetches data from APIs
Auto-Generated (Don't Edit)
public/- Built website (generated by Zola)data/- API data cache (generated by fetch_data.py)content/projects.md- Generated from GitHub reposcontent/publications.md- Generated from ORCIDcontent/blog/????-??-??-new-project-*.md- Auto-generated blog posts
These files are created automatically during deployment and are ignored by git.
Local Development
Preview Locally
# Install Zola first: https://www.getzola.org/documentation/getting-started/installation/
# Fetch latest data
pip install -r scripts/requirements.txt
python scripts/fetch_data.py
# Build and serve
zola serve
# Visit http://127.0.0.1:1111Making Changes
- Edit content in
content/folder - Modify templates in
templates/ - Update styles in
static/style.css - Test with
zola serve - Commit and push to
mainbranch - GitHub Actions deploys automatically!
GitHub Pages Setup
First Time Setup
- Go to: Settings → Pages
- Source: Deploy from a branch
- Branch: gh-pages
- Folder: / (root)
- Click Save
Requirements
- Repository must be public (for free GitHub accounts)
- GitHub Pages must be enabled in Settings
- Workflow runs successfully (check Actions tab)
Customization
Site Settings
Edit config.toml:
- Site title and description
- Base URL
- Author information
Content
Edit files in content/:
_index.md- Home pagecv.md- CV pagecontact.md- Contact pageblog/2026-02-11-welcome.md- Welcome blog post
Styling
static/style.css- Main stylesheettemplates/- HTML templates
Troubleshooting
Website Not Updating?
- Check Actions tab - look for green checkmark ✓
- If build failed, check error logs
- Verify GitHub Pages is enabled in Settings
- Clear browser cache (Ctrl+Shift+R)
Build Failing?
- Check Actions tab for error details
- Most common: Zola syntax errors in content files
- Fix the error and push again
Branch Structure
- main - Source code (you edit here)
- gh-pages - Deployed website (auto-generated, don't edit)
- copilot/ - Development branches
Advanced
Custom Domain
- Add
static/CNAMEwith your domain - Update
base_urlinconfig.toml - Configure DNS at your domain registrar
- Add custom domain in Settings → Pages
API Configuration
Edit scripts/fetch_data.py:
GITHUB_USERNAME- Your GitHub usernameORCID_ID- Your ORCID identifier- Add other data sources as needed
surveyWorkbench
Sprache: HTML
Zuletzt aktualisiert: 2026-06-30
README anzeigen
Survey Workbench v2.1
A comprehensive tool for managing participant folders and extracting survey data from PDF questionnaires with automated field mapping and Excel integration.
Overview
Survey Workbench is designed to streamline the workflow of:
- Generating standardized participant folders with multiple questionnaire types
- Prefilling PDF forms with default values before participant handover
- Extracting completed survey data directly into Excel/CSV masterfiles
- Managing field mappings for user-friendly column names
Key Features
- Dynamic Questionnaire Configuration — Support unlimited questionnaire types per participant
- Auto-Generated Field Mapping — Automatically scans form templates and generates field mappings
- Non-Participant-Specific Columns — Multiple participants in a single masterfile without naming conflicts
- Grouped Checkboxes — Automatically collapses checkbox arrays (Check1_1, Check1_2, etc.) into single columns
- Prefill Dialog — Populate PDF form fields before participant handover with a graphical interface
- Multiple Export Formats — Support for CSV, XLS, and XLSX masterfiles
- Configuration Persistence — Save and load complete project configurations including field mappings
- Type-Safe Codebase — Full Python type hints for reliability and IDE support
System Requirements
- OS: Windows 10/11
- Python: 3.10+ (if running from source)
- Excel: Microsoft Excel (for XLS/XLSX export)
- RAM: 2GB minimum
- Storage: Sufficient space for participant folders and questionnaire PDFs
Installation
Using the Executable
- Download
survey_workbench_v2.1.exefrom thedist/folder - Run the executable — no installation needed
- The first launch will create a
config.inifile in the same directory
From Source
# Clone or extract the repository
cd survey_workbench - v2.1
# Install dependencies
pip install -r requirements.txt
# Run the application
python survey_workbench_v2.1.py
Dependencies:
- PyQt5 (GUI framework)
- xlwings (Excel integration)
- pypdf (PDF processing)
Quick Start
1. Generate Participant Folders
-
Configure Questionnaires
- Enter participant ID (e.g.,
P_001) - Select target folder for participant directories
- Add questionnaire types (name, PDF template, copy count)
- Enter participant ID (e.g.,
-
Generate
- Click "Generate Participant Folder"
- System creates organized folder structure
- Prefill dialog appears for optional value population
-
Prefill (Optional)
- Edit field values in the dialog
- Check/uncheck options for checkboxes
- Click "Confirm" to save prefill values to PDFs
2. Extract Survey Data
-
Configure Extraction
- Select source folder (contains completed participant folders)
- Select masterfile (CSV, XLS, or XLSX)
- Enter participant ID to extract
-
Extract
- Click "Extract Data to Masterfile"
- System automatically detects file format
- Data is appended as new row with proper field mapping
-
Verify
- Check masterfile for new row with mapped column names
- Checkpoint values appear as option numbers (1, 2, 3, etc.)
3. Manage Field Mappings
-
Edit Field Mapping
- Click "File → Save/Edit Field Mapping"
- System auto-scans form templates and shows detected fields
- Edit "Name" column to set user-friendly column headers
- Click "Save" to persist to project configuration
-
Save Configuration
- Click "File → Save Configuration"
- Enter configuration name
- Settings, questionnaires, AND field mapping are saved together
-
Load Configuration
- Click "File → Load Configuration" (from menu)
- Select previous configuration
- Everything restores: questionnaires, paths, AND field mapping
Data Organization
Column Naming
- Previous (v2.0):
P_001_Demographics_Age(participant-specific) - Current (v2.1):
Age(mapped friendly name) orDemographics_Age(system name) - Benefit: Multiple participants can coexist in one masterfile
Checkbox Handling
Checkboxes in PDFs use naming convention:
Check1_1,Check1_2,Check1_3(individual options)- Grouped as:
Check1(single column) - Value: option number that was checked (e.g.,
1,2, or3) - System translates to/from PDF values (
/Yes,/Off,/On) automatically
Example Data Row
| participant_id | Age | Gender | Mental_Demand | Score |
|---|---|---|---|---|
| P_001 | 28 | 1 | 7.5 | 68 |
| P_002 | 35 | 2 | 5.2 | 72 |
Configuration Files
config.ini
Stores all project configurations. Each [section] is a named configuration:
[MyProject]
target_path = C:\Studies\MyProject\Participants
source_path = C:\Studies\MyProject\Completed
excel_path = C:\Studies\MyProject\masterfile.xlsx
quest_count = 3
quest_0_name = Demographics
quest_0_path = C:\Templates\demographics.pdf
quest_0_count = 1
field_mapping_json = {"Demographics_Age": "Age", "Check1": "Gender"}
Each configuration saves:
- Questionnaire setup (names, paths, counts)
- Source/target/excel paths
- Field mapping (system names → friendly names)
Workflow Example
1. Create configuration "MyStudy_2026"
├─ Add questionnaires: demographics.pdf, survey-tlx.pdf
├─ Set target folder: C:\Studies\MyStudy\Participants
└─ Save configuration
2. Generate participant folders
├─ ID: P_001
├─ Creates: P_001/
│ ├─ P_001_demographics.pdf
│ └─ P_001_survey-tlx.pdf
└─ Prefill optional values
3. Participant completes questionnaires
└─ Returns completed PDFs
4. Edit field mapping (optional)
├─ "Demographics_Age" → "Age"
├─ "Check1" → "Gender"
└─ Save to configuration
5. Extract data
├─ Source: C:\Studies\MyStudy\Completed\P_001/
├─ Masterfile: C:\Studies\MyStudy\masterfile.xlsx
└─ Appends row with mapped column namesMenu Reference
File Menu
- Save Configuration — Save current setup with field mappings
- Load Configuration — Load previous project setup
- Delete Configuration — Remove saved configuration
- Save/Edit Field Mapping — Edit field-to-column-name mappings (auto-scans templates)
- Exit — Close the application
Generation Section
- Generate Participant Folder — Create folders with questionnaires and optional prefill
Extraction Section
- Extract Data to Masterfile — Append completed survey data to masterfile
Technical Details
Version: 2.1 (June 2026)
Changes from v2.0:
- Field mapping auto-generated from form templates
- Non-participant-specific column names
- Integrated field mapping with project configuration
- Removed redundant load/delete field mapping operations
- Enhanced type hints and code quality
Technology Stack
- Python: 3.13.7
- GUI: PyQt5 5.15.11
- PDF Processing: pypdf 4.0+
- Excel Integration: xlwings 0.33.20
- Packaging: PyInstaller 6.14.2
Supported File Formats
- PDF: Standard AcroForm fields (text, checkboxes, radio buttons)
- CSV: UTF-8 encoded, comma-delimited
- Excel: XLS (2003), XLSX (2007+)
Troubleshooting
"No fields found in configured templates"
- Verify template PDF paths are correct
- Ensure PDFs have form fields (AcroForm structure)
- Check file permissions
"Field mapping does not match configured forms"
- Forms have changed since configuration was created
- Click "Save/Edit Field Mapping" to regenerate from current templates
- Review and adjust friendly names, then save
Excel file not updating
- Ensure masterfile is not open in Excel
- Verify path is correct and file is accessible
- Check that sheet named "Data" exists (or first sheet is used)
Checkbox values not appearing
- Verify checkbox field names follow pattern:
Check{N}_{option} - Ensure option numbers are sequential (1, 2, 3...)
- Check that at least one checkbox is checked in the PDF
Documentation
- USER_MANUAL.pdf — Comprehensive user guide with screenshots and step-by-step instructions
- README.md — This file, quick reference and overview
Support
For issues, feature requests, or questions:
- Check USER_MANUAL.pdf for detailed guidance
- Review configuration in config.ini
- Verify file paths and permissions
- Consult "Troubleshooting" section above
License
This project is licensed under the MIT License — See LICENSE file for details.
Copyright © 2026 Cagatay Özcan Jagiello Gutt (Original Creator and Developer)
The MIT License permits free use, modification, and distribution while preserving this copyright attribution. Your creative work and engineering effort are permanently credited as the original creator of this tool.
Survey Workbench v2.1 | June 2026 | Open-Source PDF Form Extraction Tool
continueAgent
Sprache: Python
Zuletzt aktualisiert: 2026-06-25
README anzeigen
ContinueAutomation Agent
Intelligent token optimization agent for VS Code Continue extension. Automatically selects between Mistral and Claude based on question complexity, simplifies queries, and tracks costs in real-time.
Features
- Smart Model Selection: Analyzes question complexity and auto-switches between Mistral (simple) and Claude (complex)
- Query Simplification: Reduces token usage by simplifying questions while preserving meaning
- Cost Tracking: Real-time token and cost estimation before and after each query
- Auto-Install: One-command installation that runs as background service on startup
- Prompt Caching: Caches frequently used prompts for efficiency
Installation from GitHub
git clone https://github.com/yourusername/continueAgent.git
cd continueAgent
python install_service.py
This will:
- Install dependencies (Flask, PyYAML, requests)
- Create startup service (auto-starts on login)
- Start API server on http://localhost:5001
Configure Continue
Add to ~/.continue/config.yaml:
- name: ContinueAutomation
provider: openai
model: continue-automation
apiBase: http://localhost:5001/v1
apiKey: dummy
Restart VS Code and select "ContinueAutomation" from the model dropdown.
How It Works
- Initial Assessment: Analyzes question, decides model, simplifies query, shows estimated cost
- Model Call: Sends simplified question to chosen model (Mistral/Claude)
- Final Report: Shows actual token usage and cost
License
MIT License - see LICENSE file
tVNS_deviceplan
Sprache: Unbekannt
Zuletzt aktualisiert: 2026-06-10
README anzeigen
tVNS_deviceplan
GitRef
Sprache: Python
Zuletzt aktualisiert: 2026-04-15
README anzeigen
GitRef
A git-based reference manager — like Zotero, but your library is a plain git repo.
Install
pip (recommended)
pip install gitrefStandalone binary
Download from the Releases page:
| Platform | Asset | Install |
|---|---|---|
| Windows | gitref.exe | Move to a folder on your PATH, or run: irm https://raw.githubusercontent.com/CGutt-hub/gitref/main/install.ps1 | iex |
| Linux | GitRef-x86_64.AppImage or gitref | chmod +x and move to /usr/local/bin/, or run: curl -fsSL https://raw.githubusercontent.com/CGutt-hub/gitref/main/install.sh | bash |
From source
pip install git+https://github.com/CGutt-hub/gitref.gitFeatures
- DOI / arXiv / ISBN lookup – paste an identifier, metadata is fetched automatically
- PDF download – auto-downloads from arXiv or DOI resolvers
- BibTeX store – all references in
.resources.bib(human-readable, diffable, standard) - Zip archive – PDFs stored compactly in
.resources.zipfor git efficiency - File watcher –
gitref watchextracts all PDFs, auto-adds dropped files, auto-repacks on close - Clickable links –
resources.mdhas direct PDF links when watcher is running - Git sync – auto-commit + pull + push to keep your library in sync across machines
- Terminal UI – interactive browser with search, detail view, tagging
- Browser extension – one-click save from any paper page (like Zotero Connector)
Quick start
# Initialise a new library
gitref init
# Add a paper by DOI
gitref add "10.1038/s41586-024-07487-w"
# Add by arXiv ID
gitref add "2401.12345"
# Start watcher (extracts PDFs, watches for new/modified files)
gitref watch
# Open interactive TUI
gitref browse
# Sync with remote
gitref syncWorkflow
DAU-friendly PDF access — no need to use gitref open / gitref close:
- Run
gitref watch— this extracts all PDFs from the archive - Open
resources/resources.mdand click any 📎 link to read a paper - Drop a PDF into
resources/and it's auto-added to the library - Close your PDF viewer (or just hit Ctrl+C on the watcher) — files are repacked
Browser extension
GitRef includes a Chrome/Edge extension (Manifest V3) that works like the Zotero Connector — click the toolbar icon on any paper page to save it to your library.
Setup
- Start the local server:
gitref serve - Load the extension:
- Open
chrome://extensions(oredge://extensions) - Enable Developer mode
- Click Load unpacked and select the
extension/folder from this repo
- Open
- Navigate to any paper page and click the GitRef icon — the reference (and PDF if available) is saved automatically.
The extension auto-detects DOIs, arXiv IDs, and citation metadata from the page, just like Zotero Connector.
Library structure
~/GitRef/
├── .git/
├── .github/workflows/ # auto-regenerates resources.md on push
├── README.md
└── resources/
├── .resources.bib # BibTeX metadata (source of truth)
├── .resources.zip # all PDFs, compressed
└── resources.md # browsable table with clickable PDF linksCommands
Run gitref --help for full usage. Key commands:
| Command | Description |
|---|---|
gitref init | Create a new library (git repo + bib) |
gitref add <id> | Add by DOI, arXiv ID, ISBN, or URL |
gitref search <q> | Search title, authors, tags |
gitref list | Print all entries |
gitref browse | Interactive TUI |
gitref watch | Extract all PDFs, auto-add/repack on changes |
gitref open <key> | Extract a single PDF for reading |
gitref close <key> | Repack a PDF into archive |
gitref compact | Pack all loose PDFs into archive |
gitref sync | Git add + commit + pull + push |
gitref serve | Start server for browser extension (port 7342) |
gitref export | Export library as RIS |
Use -l <path> to specify a custom library location (default: ~/GitRef).
labourAIVolt
Sprache: Python
Zuletzt aktualisiert: 2026-03-18
README anzeigen
labourAIVolt: AI & Human Labour Displacement Analysis for Volt
This repository hosts a Nextflow + Python analysis pipeline that automatically fetches current labour-market data from the World Bank public API and quantifies AI-driven labour displacement across the six Volt EU countries with the most active chapters — Germany, France, the Netherlands, Belgium, Italy, and Spain.
The aim is to give Volt Europa and its national chapters an evidence base for labour-market and technology policy: which sectors are shedding jobs fastest as AI and automation accelerate, which countries are most exposed, and how does digital readiness moderate that exposure?
The pipeline is architecturally modelled after the EmotiView project and extends the AnalysisToolbox modular Nextflow framework for scalable, reproducible analysis with automatic result synchronisation.
| Platform | Role | Contents |
|---|---|---|
| GitHub | Technical implementation | Pipeline, scripts, results |
| World Bank API | Data source | Live labour-market indicators |
Research Background
The labour-market impact of AI and automation is one of the defining policy challenges of the 2020s. Early projections (Frey & Osborne, 2013) estimated that up to 47 % of US jobs faced high computerisation risk; subsequent analyses have moderated that figure while broadening it to task-level disruption rather than wholesale job destruction. What is clear is that the pace and sector distribution of displacement vary substantially across countries depending on industrial structure, education levels, and digital infrastructure.
For a pan-European political movement like Volt, the relevant questions are:
- Which EU sectors show the clearest employment decline correlated with automation?
- Are Volt's home countries converging toward or diverging from each other on displacement pressure?
- Does a country's digital readiness (internet penetration, high-tech exports) buffer it against displacement?
- Where should Volt's labour policy — reskilling funds, working-time reform, Universal Basic Income pilots — be concentrated first?
This pipeline operationalises those questions with reproducible, automatically updated data.
Core Research Questions & Hypotheses
-
Sector displacement ordering: Industry and agriculture will show larger negative employment-share trends than services, consistent with higher Frey & Osborne automation risk for routine physical/cognitive tasks.
-
Cross-country heterogeneity: Countries with larger manufacturing sectors (Germany, Italy) will exhibit higher AI Displacement Pressure Index (ADPI) than service-dominant economies (Netherlands, Belgium).
-
Digitalization buffer: Countries with higher Digitalization Readiness Scores (internet penetration + high-tech export share) will show lower net Vulnerability Scores, suggesting that digital transformation simultaneously creates displacement and provides adaptive capacity.
-
Temporal acceleration: Employment-share trends will steepen post-2018 as AI adoption accelerates across all three sectors, visible as a structural break in the time-series.
Analysis Pipeline
The pipeline is built on the
AnalysisToolbox — a modular Nextflow
framework for scalable, reproducible data processing with automatic result synchronisation.
The LAV-specific pipeline in LAV_analysis/ extends this framework to:
- Discover country datasets and create per-country output directories (L1).
- Fetch and parse live labour-market time-series from the World Bank API.
- Perform standardised normalisation and cleaning.
- Extract displacement signals and automation-risk-weighted scores per sector.
- Fit time-series trend models to all available indicators.
- Aggregate across countries into cross-country rankings and Volt policy metrics (L2).
Configuration is managed via LAV_analysis/LAV_parameters.config.
See the AnalysisToolbox documentation
for framework details.
Pipeline steps
L1 (per country)
┌─────────────────────────────────────────────────────────────────────┐
│ api_reader Fetch 13 World Bank indicators (2010–2023) │
│ ↓ │
│ normalizing_processor Pivot long→wide; sort; deduplicate │
│ ↓ ↘ │
│ displacement_analyzer trend_analyzer │
│ sector scores × OLS slope + p-value + R² │
│ Frey & Osborne risk per indicator │
└─────────────────────────────────────────────────────────────────────┘
↓ collect all countries
L2 (cross-country)
┌─────────────────────────────────────────────────────────────────────┐
│ volt_report_analyzer Displacement ranking · ADPI · DRS │
│ Vulnerability Score · policy metrics │
└─────────────────────────────────────────────────────────────────────┘Displacement model
Each broad employment sector receives a displacement score:
displacement_score = displacement_signal × automation_risk| Term | Definition |
|---|---|
displacement_signal | Normalised negative employment-share trend: max(0, −slope / mean_level). A sector losing share faster relative to its baseline scores higher. |
automation_risk | Sector-level probability of computerisation from Frey & Osborne (2013): agriculture 0.82, industry 0.79, services 0.63. |
displacement_score | Composite: high score = fast employment decline and high intrinsic automation susceptibility. |
Group-level metrics (L2)
| Metric | Definition |
|---|---|
| ADPI (AI Displacement Pressure Index) | Mean displacement_score across all sectors for a country. |
| DRS (Digitalization Readiness Score) | Normalised mean of internet-user percentage and high-tech export share (latest year). |
| Vulnerability Score | ADPI / (DRS + ε) — high ADPI and low digital readiness = most vulnerable. |
Data indicators (World Bank API, no key required)
| Column | World Bank code | Description |
|---|---|---|
employment_agriculture_pct | SL.AGR.EMPL.ZS | Employment in agriculture (% total) |
employment_industry_pct | SL.IND.EMPL.ZS | Employment in industry (% total) |
employment_services_pct | SL.SRV.EMPL.ZS | Employment in services (% total) |
unemployment_rate | SL.UEM.TOTL.ZS | Unemployment (% labour force) |
youth_unemployment_rate | SL.UEM.1524.ZS | Youth unemployment (%) |
employment_to_pop_ratio | SL.EMP.TOTL.SP.ZS | Employment-to-population ratio |
wage_salary_workers_pct | SL.EMP.WORK.ZS | Wage & salaried workers (%) |
internet_users_pct | IT.NET.USER.ZS | Internet users (% population) |
gdp_per_capita_usd | NY.GDP.PCAP.CD | GDP per capita (current USD) |
gdp_growth_annual_pct | NY.GDP.MKTP.KD.ZG | GDP growth (annual %) |
high_tech_exports_pct_mfg | TX.VAL.TECH.MF.ZS | High-tech exports (% manufactured exports) |
ict_goods_exports_pct | TX.VAL.ICTG.ZS.UN | ICT goods exports (% total goods exports) |
labor_force_total | SL.TLF.TOTL.IN | Total labour force |
Repository Structure
labourAIVolt/
│
├── LAV_analysis/ Nextflow pipeline (mirrors EV_analysis/)
│ ├── LAV_pipeline.nf Main workflow orchestration
│ ├── LAV_modules.nf IOInterface alias declarations
│ └── LAV_parameters.config All pipeline parameters & script paths
│
├── LAV_data/ Per-country input configs (mirrors rawData/)
│ ├── LAV_001/ LAV_001_config.json Germany (Volt Deutschland)
│ ├── LAV_002/ LAV_002_config.json France (Volt France)
│ ├── LAV_003/ LAV_003_config.json Netherlands (Volt Nederland)
│ ├── LAV_004/ LAV_004_config.json Belgium (Volt Belgium)
│ ├── LAV_005/ LAV_005_config.json Italy (Volt Italia)
│ └── LAV_006/ LAV_006_config.json Spain (Volt España)
│
├── LAV_results/ Pipeline outputs (mirrors EV_results/)
│ ├── .bin/ Shared infrastructure (logs, HTML archive)
│ ├── LAV_l1/ First-level: per-country results
│ │ ├── LAV_001/
│ │ │ ├── plots/ Parquet output copies for QC
│ │ │ ├── LAV_001_api_raw.parquet
│ │ │ ├── LAV_001_normalized.parquet
│ │ │ ├── LAV_001_displacement.parquet
│ │ │ ├── LAV_001_trends.parquet
│ │ │ └── LAV_001.log.parquet Live execution log
│ │ └── LAV_002/ … LAV_006/
│ └── LAV_l2/ Second-level: cross-country group results
│ ├── LAV_volt_report.parquet
│ ├── LAV_displacement_summary.parquet
│ └── LAV_trends_summary.parquet
│
├── Python/ Analysis scripts (no Nextflow dependency)
│ ├── lav_run.py Standalone orchestrator (used by CI)
│ ├── requirements.txt
│ ├── readers/
│ │ └── api_reader.py Fetches World Bank labour-market data
│ ├── processors/
│ │ └── normalizing_processor.py Long→wide pivot, clean, sort
│ └── analyzers/
│ ├── displacement_analyzer.py AI displacement scores (Frey & Osborne)
│ ├── trend_analyzer.py OLS time-series trends per indicator
│ └── volt_report_analyzer.py Cross-country Volt policy synthesis
│
└── .github/workflows/
└── lav_analysis.yml GitHub Actions CI (weekly + on push)
Running the Analysis
Option A — GitHub Actions (recommended — no local setup needed)
The workflow in .github/workflows/lav_analysis.yml runs automatically:
| Trigger | When |
|---|---|
| Scheduled | Every Monday at 06:00 UTC (pulls the latest World Bank data) |
| On push | Any change to LAV_data/** or Python/** on main |
| Manual | Actions tab → LAV Labour-AI-Volt Analysis → Run workflow |
Results are:
- Uploaded as a downloadable artifact (
lav-results-<run-number>) for 90 days. - Committed back to
LAV_results/in the repository so outputs are versioned alongside the code.
No API keys, secrets, or local software are required.
Option B — Standalone Python (local, no Nextflow)
Use this for quick local runs or debugging individual scripts.
# 1. Clone the repository
git clone https://github.com/CGutt-hub/labourAIVolt.git
cd labourAIVolt
# 2. Install Python dependencies
pip install -r Python/requirements.txt
# 3. Run the full pipeline
python Python/lav_run.py
# Optional: override data/output directories
python Python/lav_run.py --data-dir LAV_data --output-dir LAV_results
Results are written to LAV_results/LAV_l1/<id>/ (per country) and
LAV_results/LAV_l2/ (group synthesis).
Option C — Full Nextflow Pipeline (local, requires AnalysisToolbox)
Use this for full pipeline tracing, parallel execution, and integration with the AnalysisToolbox interactive HTML archive.
Prerequisites: Java ≥ 11, Nextflow
# 1. Clone both repos as siblings
git clone https://github.com/CGutt-hub/labourAIVolt.git
git clone https://github.com/CGutt-hub/AnalysisToolbox.git
# Your directory should now look like:
# parent/
# ├── AnalysisToolbox/
# └── labourAIVolt/
# 2. Install Python dependencies
cd labourAIVolt
pip install -r Python/requirements.txt
# 3. Adjust python_exe in LAV_parameters.config if needed
# (default: 'python3')
# 4. Launch the pipeline from the LAV_analysis/ directory
cd LAV_analysis
nextflow run LAV_pipeline.nf -c LAV_parameters.config
The Nextflow pipeline adds on top of the standalone runner:
- Parallel per-country execution
- Full Nextflow trace (
LAV_results/.bin/pipeline_trace.txt) - Interactive HTML result archive (via AnalysisToolbox
interactive_plotter) - Automatic git commit + push of results after each country completes
Output Files Reference
Per-country (L1) — LAV_results/LAV_l1/LAV_XXX/
| File | Description |
|---|---|
LAV_XXX_api_raw.parquet | Raw long-format data as returned by the World Bank API. Columns: participant_id, country, iso3, source, indicator, indicator_code, year, value. |
LAV_XXX_normalized.parquet | Wide-format time-series. One row per year, one column per indicator. Ready for analysis scripts. |
LAV_XXX_displacement.parquet | Per-sector displacement scores. Key columns: sector, employment_mean_pct, trend_slope_pp_per_yr, trend_significant, automation_risk_frey_osborne, displacement_score. |
LAV_XXX_trends.parquet | OLS trend results for every indicator. Key columns: indicator, trend_slope, trend_p_value, trend_r_squared, trend_significant, total_change_pct. |
LAV_XXX.log.parquet | Live pipeline execution log (Nextflow mode only). |
Group-level (L2) — LAV_results/LAV_l2/
| File | Description |
|---|---|
LAV_volt_report.parquet | Full combined table (displacement + policy metrics for all countries). |
LAV_displacement_summary.parquet | Cross-country displacement ranking per sector, with EU-wide mean, std, and per-country rank. |
LAV_trends_summary.parquet | EU-wide mean slope and significance counts for key indicators across all countries. |
Adding a New Country
- Create a new directory:
LAV_data/LAV_007/ - Add a config file
LAV_data/LAV_007/LAV_007_config.json:
{
"participant_id": "LAV_007",
"country": "Portugal",
"iso3": "PRT",
"iso2": "PT",
"year_start": 2010,
"year_end": 2025,
"volt_chapter": "Volt Portugal",
"population_millions": 10.3,
"eu_member": true,
"notes": "Optional notes about the country context"
}
- Push the file — the GitHub Action will pick it up automatically on the next run.
Project Status
Active development. Data fetching, pipeline, and group analysis are operational. Planned additions: visualisation layer, structural-break detection (2018 AI inflection point), and integration with OECD employment-by-occupation microdata for finer-grained occupational risk scoring.
References
- Frey, C. B., & Osborne, M. A. (2013). The Future of Employment: How Susceptible Are Jobs to Computerisation? Oxford Martin School Working Paper.
- World Bank Open Data. https://data.worldbank.org
- Acemoglu, D., & Restrepo, P. (2020). Robots and Employment: Evidence from Europe. American Economic Review, 110(6), 2188–2220.
- Autor, D. (2015). Why Are There Still So Many Jobs? Journal of Economic Perspectives, 29(3), 3–30.
Contributors
| Name | Role | Contact |
|---|---|---|
| Cagatay Özcan Jagiello Gutt | Principal Investigator |
paperFinder
Sprache: Python
Zuletzt aktualisiert: 2026-02-18
README anzeigen
Entwicklungsphilosophie
Aller Code wird mit dem Engagement für offene und transparente Wissenschaft entwickelt. Werkzeuge, Pipelines und Analysecode werden verfügbar gemacht, um Reproduzierbarkeit und kollaborativen Wissensfortschritt zu unterstützen.