EN | DE

Code Projects & Repositories

Active development projects tracked via GitHub. My open backoffice for collaborative science.


EmotiView

Language: Jupyter Notebook ⭐ 1
Last updated: 2026-08-18

View README

EmotiView: Neural-Autonomic Synchrony and Embodied Integration

This repository accompanies ongoing research investigating the dynamic interplay between neural activity and autonomic nervous system responses during emotional experiences. Here you'll find the research article, presentations, analysis pipeline, and results—updated in real-time as the project progresses.

Principal Investigator: Cagatay Özcan Jagiello Gutt

PlatformRoleContents
OSFResearch outputArticle, documentation
GitHubTechnical implementationAnalysis pipeline, results, presentations, proposal

Research Abstract

Emotional states are fundamentally embodied, emerging from the dynamic interplay between central neural processing and peripheral physiological adjustments orchestrated by the autonomic nervous system (ANS). While ANS outputs like heart rate variability (HRV) and electrodermal activity (EDA) reflect emotional arousal and valence, understanding the precise temporal coordination between brain activity and these peripheral signals is crucial for elucidating brain-body interactions. This study investigates neural-autonomic phase synchrony during the conscious processing of distinct emotional states (positive, negative, neutral) by quantifying the temporal alignment between cortical and physiological rhythms.

We employ a multimodal approach, simultaneously recording high-temporal-resolution electroencephalography (EEG), electrocardiography (ECG) for HRV analysis (specifically Root Mean Square of Successive Differences, RMSSD), EDA, and functional near-infrared spectroscopy (fNIRS) while participants view validated emotional video clips. Our primary analysis quantifies the Phase Locking Value (PLV) between frontal EEG oscillations (Alpha, Beta bands) and continuous signals derived from HRV (reflecting parasympathetic influence) and phasic EDA (reflecting sympathetic influence). EEG channel selection for PLV analysis is informed by task-related hemodynamic activity measured via fNIRS to focus on functionally relevant cortical areas.

We hypothesize that PLV, indicating brain-body temporal integration, will be significantly modulated by emotional content compared to neutral conditions. We further expect synchrony strength to correlate with subjective arousal ratings. By examining the phase synchrony between brain signals and ANS-mediated physiological outputs, this research provides novel insights into the dynamic, embodied mechanisms underlying emotional experience. Understanding this temporal binding is critical for models of psychophysiological function and may inform assessments of cognitive load or stress regulation capacity.

Core Research Aims & Hypotheses

This project seeks to understand how the brain and body coordinate during emotional processing, focusing on neural-autonomic phase synchrony. Key hypotheses include:

  1. Emotional Modulation of Synchrony: Neural-autonomic synchrony (Phase Locking Value - PLV) will be enhanced during the processing of positive and negative emotional stimuli compared to neutral stimuli, for both brain-heart (EEG-HRV) and brain-sudomotor (EEG-EDA) coupling.
  2. Synchrony and Subjective Arousal: The magnitude of neural-autonomic synchrony will positively correlate with subjective ratings of emotional arousal during emotional conditions.
  3. Baseline Vagal Tone and Task-Related Synchrony: Individual differences in baseline parasympathetic regulation (resting-state RMSSD) will be associated with the degree of EEG-HRV synchrony during negative emotional stimuli.
  4. Frontal Asymmetry and Branch-Specific Synchrony: The direction of prefrontal cortical asymmetry (Frontal Asymmetry Index - FAI) will be differentially associated with the strength of phase synchrony involving distinct autonomic branches (EEG-HRV vs. EEG-EDA).

For a comprehensive understanding of the theoretical background, detailed methodology, and specific work packages, please refer to the full proposal document.

Methodology Overview

A multimodal experimental design is employed, involving:

  • Stimuli: Standardized, emotionally evocative video clips (positive, negative, neutral) from the E-MOVIE database.
  • Participants: Healthy young adults, screened for relevant criteria.
  • Data Acquisition: Simultaneous recording of:
    • Electroencephalography (EEG): To measure prefrontal neural dynamics.
    • Functional Near-Infrared Spectroscopy (fNIRS): To localize hemodynamic activity in prefrontal and parietal regions, informing EEG channel selection.
    • Electrocardiography (ECG): For Heart Rate Variability (HRV) analysis.
    • Electrodermal Activity (EDA): To measure sympathetic nervous system activity.
  • Subjective Measures: Self-Assessment Manikin (SAM) for valence and arousal, Positive and Negative Affect Schedule (PANAS), and Behavioural Inhibition/Approach System (BIS/BAS) scales.

Repository Contents

Research Output (OSF)

  • Article: (Coming soon) The research article summarizing findings and contributions.

Technical Implementation (GitHub)

  • EV_results/: Processed data, analysis metrics, and visualizations.
  • EV_analysis/: The Nextflow-based analysis pipeline with Python modules.
  • EV_presentation/: Slides and presentation materials.
  • EV_proposal/: The original research proposal with methodology and analysis plan.

Analysis Pipeline

The analysis pipeline is built on the AnalysisToolbox—a modular Nextflow framework for scalable, reproducible data processing with automatic result synchronization. The EmotiView-specific pipeline in EV_analysis/ extends this framework to:

  • Load and parse multi-modal raw data (EEG, fNIRS, ECG, EDA, questionnaires).
  • Perform standardized preprocessing steps specific to each physiological modality.
  • Extract key features and metrics (e.g., EEG power, FAI, RMSSD, fNIRS ROI activation, PLV).
  • Generate participant-level results and aggregated summaries.

Configuration is managed via EV_analysis/EV_parameters.config. See the AnalysisToolbox documentation for framework details.

Project Status

Data collection and Thesis writing.

Contributors

NameRoleContact
Cagatay Özcan Jagiello GuttPrincipal InvestigatorORCID
Ben GopinTechnical AssistantEmail
Gerrit JostlerTechnical AssistantEmail

View on GitHub →


AnalysisToolbox

Language: Python
Last updated: 2026-09-02

View README

AnalysisToolbox

A modular framework for automated data processing and statistical analysis pipelines. Built on Nextflow for scalable, reproducible workflows with automatic result synchronization.

Overview

The AnalysisToolbox provides infrastructure for building data processing pipelines that:

  • Process multiple datasets in parallel with automatic participant discovery
  • Handle diverse data types through a generic reader/processor/analyzer architecture
  • Track progress via per-participant logging visible live in the web UI
  • Recover gracefully from failures without losing completed work

The framework is domain-agnostic — modules follow simple input/output conventions (Parquet/FIF files) and can implement any processing logic.

Repository Structure

AnalysisToolbox/
├── gitatbx/               # pip-installable package (installed via pip install GitAtbx)
│   ├── bin/               # workflow_wrapper.nf, log_to_parquet.py, nextflow.config, ...
│   ├── modules/           # analyzers/, processors/, readers/, utils/
│   ├── templates/         # workflow_template.nf, modules_template.nf, parameters_template.config
│   └── utils/             # serve_html.ps1, reinject.sh, result_collector.py
├── pyproject.toml
└── README.md

On first run gitatbx creates a symlink ~/Documents/GitAtbxModules → <site-packages>/gitatbx/modules/. The symlink always reflects the live installed version — upgrading via pip install --upgrade GitAtbx automatically shows updated modules through the same symlink path.

Prerequisites

Java Runtime (required by Nextflow)

sudo apt update && sudo apt install default-jre
java -version

Nextflow

curl -s https://get.nextflow.io | bash
sudo mv nextflow /usr/local/bin/
nextflow -version

Install

pip install GitAtbx

Python dependencies (numpy, scipy, polars, mne, neurokit2, …) are installed automatically.

On first run gitatbx creates a symlink ~/Documents/GitAtbxModules → <site-packages>/gitatbx/modules/ and saves the path to ~/.gitatbx_config.

Usage

Commands

CommandWhat it does
gitatbx init <dir>Scaffold a new analysis project
gitatbx run [pattern]Find and run a pipeline by project name pattern
gitatbx serve [--dir DIR] [--port PORT]Serve results HTML locally in browser
gitatbx reinject <PID> [options]Reinject a corrected output for one participant
gitatbx move <dest>Move the deployed modules folder to a new location
gitatbx config showPrint current configuration (~/.gitatbx_config)

gitatbx init

Prompts for project name, raw data directory, Python executable, toolbox path, git author identity, and an optional GitHub remote URL for the results repo. Then creates:

<dir>/
├── {name}_analysis/
│   ├── {name}_pipeline.nf       ← edit your workflow here
│   ├── {name}_modules.nf        ← add IOInterface includes here
│   └── {name}_parameters.config ← pre-filled paths, params, and git identity
└── {name}_results/
    └── .git/                    ← initialised + remote added (if URL provided)

Git author name and email default to your global git config values if already set. The remote URL is validated immediately with git ls-remote — if authentication fails (e.g. SSH key not yet added to GitHub), a warning is printed with a link to the GitHub SSH setup guide.

The pipeline uses the stamped params.git_user_name / params.git_user_email as the commit author for all automatic result syncs.

gitatbx run

GitAtbx searches the entire accessible filesystem (home directory and all drives on Windows) for a directory named (name)_analysis containing a *_pipeline.nf, then runs it automatically. No need to cd anywhere.

gitatbx run (name) --resume   # continue a previous run
gitatbx run               # no pattern: use current directory

Found paths are cached in ~/.gitatbx_config so subsequent calls are instant.

gitatbx serve

Starts a local HTTP server to browse results HTML generated by the pipeline.

gitatbx serve --dir ../EV_results --port 8080

gitatbx reinject

Places a corrected parquet into corrections/<script_name>/, marks the participant for replay, invalidates relevant Nextflow cache entries, and resumes the pipeline for that participant only.

gitatbx reinject EV_002 --corrected-file fixed.parquet --script-name filtering_processor

gitatbx move

To move the symlink to a custom location:

gitatbx move /mnt/d/repoShaggy/GitAtbxModules

gitatbx moves the folder and updates ~/.gitatbx_config automatically. gitatbx init will then default toolbox_dir to that path when scaffolding new projects.

Key Components

bin/workflow_wrapper.nf

Discovers participant directories, manages per-participant output folders, runs log_to_parquet.py and interactive_plotter.py automatically, and handles per-participant git sync on completion.

IOInterface

Generic Nextflow process that runs any script (reader/processor/analyzer) with automatic logging. Every module in modules/ is called through IOInterface.

Modules

The modules/ folder is a curated but non-exhaustive starting collection of commonly useful scripts. The first time gitatbx is run it mirrors the full collection to ~/Documents/GitAtbxModules (Windows) or ~/GitAtbxModules (Linux/macOS) automatically. You can move that folder anywhere with gitatbx move. The local copy is yours to extend: add domain-specific scripts, modify existing ones, or organise them into subfolders. Any script there can be called via IOInterface identically to the built-in modules, as long as it follows the same convention: positional CLI arguments in, Parquet or FIF outputs, non-zero exit on failure.

Authors

Cagatay Özcan Jagiello Gutt — Lead Developer ORCID: https://orcid.org/0000-0002-1774-532X

View on GitHub →


5ha99y

Language: JavaScript
Last updated: 2026-07-02

View README

Zola GitHub Pages Site - Scientific Hub

A static website built with Zola that automatically syncs content from your scientific profiles.

Website URL

https://cgutt-hub.github.io/5ha99y

How It Works

Automatic Updates

The site automatically pulls data from:

  • GitHub — Your repositories and code projects
  • ORCID — Publications and works
  • OSF — Research projects and data (if configured)

Deployment Flow

1. Push to main branch
   ↓
2. GitHub Actions runs
   ↓
3. Fetches data from APIs
   ↓
4. Builds site with Zola
   ↓
5. Deploys to gh-pages branch
   ↓
6. Website updates automatically!

Repository Structure

Source Files (Edit These)

  • config.toml - Site configuration
  • content/ - Your content (CV, contact, welcome post, etc.)
  • templates/ - HTML templates
  • static/ - Static assets (CSS, images)
  • scripts/fetch_data.py - Fetches data from APIs

Auto-Generated (Don't Edit)

  • public/ - Built website (generated by Zola)
  • data/ - API data cache (generated by fetch_data.py)
  • content/projects.md - Generated from GitHub repos
  • content/publications.md - Generated from ORCID
  • content/blog/????-??-??-new-project-*.md - Auto-generated blog posts

These files are created automatically during deployment and are ignored by git.

Local Development

Preview Locally

# Install Zola first: https://www.getzola.org/documentation/getting-started/installation/

# Fetch latest data
pip install -r scripts/requirements.txt
python scripts/fetch_data.py

# Build and serve
zola serve
# Visit http://127.0.0.1:1111

Making Changes

  1. Edit content in content/ folder
  2. Modify templates in templates/
  3. Update styles in static/style.css
  4. Test with zola serve
  5. Commit and push to main branch
  6. GitHub Actions deploys automatically!

GitHub Pages Setup

First Time Setup

  1. Go to: Settings → Pages
  2. Source: Deploy from a branch
  3. Branch: gh-pages
  4. Folder: / (root)
  5. Click Save

Requirements

  • Repository must be public (for free GitHub accounts)
  • GitHub Pages must be enabled in Settings
  • Workflow runs successfully (check Actions tab)

Customization

Site Settings

Edit config.toml:

  • Site title and description
  • Base URL
  • Author information

Content

Edit files in content/:

  • _index.md - Home page
  • cv.md - CV page
  • contact.md - Contact page
  • blog/2026-02-11-welcome.md - Welcome blog post

Styling

  • static/style.css - Main stylesheet
  • templates/ - HTML templates

Troubleshooting

Website Not Updating?

  1. Check Actions tab - look for green checkmark ✓
  2. If build failed, check error logs
  3. Verify GitHub Pages is enabled in Settings
  4. Clear browser cache (Ctrl+Shift+R)

Build Failing?

  • Check Actions tab for error details
  • Most common: Zola syntax errors in content files
  • Fix the error and push again

Branch Structure

  • main - Source code (you edit here)
  • gh-pages - Deployed website (auto-generated, don't edit)
  • copilot/ - Development branches

Advanced

Custom Domain

  1. Add static/CNAME with your domain
  2. Update base_url in config.toml
  3. Configure DNS at your domain registrar
  4. Add custom domain in Settings → Pages

API Configuration

Edit scripts/fetch_data.py:

  • GITHUB_USERNAME - Your GitHub username
  • ORCID_ID - Your ORCID identifier
  • Add other data sources as needed

View on GitHub →


surveyWorkbench

Language: HTML
Last updated: 2026-06-30

View README

Survey Workbench v2.1

A comprehensive tool for managing participant folders and extracting survey data from PDF questionnaires with automated field mapping and Excel integration.

Overview

Survey Workbench is designed to streamline the workflow of:

  1. Generating standardized participant folders with multiple questionnaire types
  2. Prefilling PDF forms with default values before participant handover
  3. Extracting completed survey data directly into Excel/CSV masterfiles
  4. Managing field mappings for user-friendly column names

Key Features

  • Dynamic Questionnaire Configuration — Support unlimited questionnaire types per participant
  • Auto-Generated Field Mapping — Automatically scans form templates and generates field mappings
  • Non-Participant-Specific Columns — Multiple participants in a single masterfile without naming conflicts
  • Grouped Checkboxes — Automatically collapses checkbox arrays (Check1_1, Check1_2, etc.) into single columns
  • Prefill Dialog — Populate PDF form fields before participant handover with a graphical interface
  • Multiple Export Formats — Support for CSV, XLS, and XLSX masterfiles
  • Configuration Persistence — Save and load complete project configurations including field mappings
  • Type-Safe Codebase — Full Python type hints for reliability and IDE support

System Requirements

  • OS: Windows 10/11
  • Python: 3.10+ (if running from source)
  • Excel: Microsoft Excel (for XLS/XLSX export)
  • RAM: 2GB minimum
  • Storage: Sufficient space for participant folders and questionnaire PDFs

Installation

Using the Executable

  1. Download survey_workbench_v2.1.exe from the dist/ folder
  2. Run the executable — no installation needed
  3. The first launch will create a config.ini file in the same directory

From Source

# Clone or extract the repository
cd survey_workbench - v2.1

# Install dependencies
pip install -r requirements.txt

# Run the application
python survey_workbench_v2.1.py

Dependencies:

  • PyQt5 (GUI framework)
  • xlwings (Excel integration)
  • pypdf (PDF processing)

Quick Start

1. Generate Participant Folders

  1. Configure Questionnaires

    • Enter participant ID (e.g., P_001)
    • Select target folder for participant directories
    • Add questionnaire types (name, PDF template, copy count)
  2. Generate

    • Click "Generate Participant Folder"
    • System creates organized folder structure
    • Prefill dialog appears for optional value population
  3. Prefill (Optional)

    • Edit field values in the dialog
    • Check/uncheck options for checkboxes
    • Click "Confirm" to save prefill values to PDFs

2. Extract Survey Data

  1. Configure Extraction

    • Select source folder (contains completed participant folders)
    • Select masterfile (CSV, XLS, or XLSX)
    • Enter participant ID to extract
  2. Extract

    • Click "Extract Data to Masterfile"
    • System automatically detects file format
    • Data is appended as new row with proper field mapping
  3. Verify

    • Check masterfile for new row with mapped column names
    • Checkpoint values appear as option numbers (1, 2, 3, etc.)

3. Manage Field Mappings

  1. Edit Field Mapping

    • Click "File → Save/Edit Field Mapping"
    • System auto-scans form templates and shows detected fields
    • Edit "Name" column to set user-friendly column headers
    • Click "Save" to persist to project configuration
  2. Save Configuration

    • Click "File → Save Configuration"
    • Enter configuration name
    • Settings, questionnaires, AND field mapping are saved together
  3. Load Configuration

    • Click "File → Load Configuration" (from menu)
    • Select previous configuration
    • Everything restores: questionnaires, paths, AND field mapping

Data Organization

Column Naming

  • Previous (v2.0): P_001_Demographics_Age (participant-specific)
  • Current (v2.1): Age (mapped friendly name) or Demographics_Age (system name)
  • Benefit: Multiple participants can coexist in one masterfile

Checkbox Handling

Checkboxes in PDFs use naming convention:

  • Check1_1, Check1_2, Check1_3 (individual options)
  • Grouped as: Check1 (single column)
  • Value: option number that was checked (e.g., 1, 2, or 3)
  • System translates to/from PDF values (/Yes, /Off, /On) automatically

Example Data Row

participant_idAgeGenderMental_DemandScore
P_0012817.568
P_0023525.272

Configuration Files

config.ini

Stores all project configurations. Each [section] is a named configuration:

[MyProject]
target_path = C:\Studies\MyProject\Participants
source_path = C:\Studies\MyProject\Completed
excel_path = C:\Studies\MyProject\masterfile.xlsx
quest_count = 3
quest_0_name = Demographics
quest_0_path = C:\Templates\demographics.pdf
quest_0_count = 1
field_mapping_json = {"Demographics_Age": "Age", "Check1": "Gender"}

Each configuration saves:

  • Questionnaire setup (names, paths, counts)
  • Source/target/excel paths
  • Field mapping (system names → friendly names)

Workflow Example

1. Create configuration "MyStudy_2026"
   ├─ Add questionnaires: demographics.pdf, survey-tlx.pdf
   ├─ Set target folder: C:\Studies\MyStudy\Participants
   └─ Save configuration

2. Generate participant folders
   ├─ ID: P_001
   ├─ Creates: P_001/
   │   ├─ P_001_demographics.pdf
   │   └─ P_001_survey-tlx.pdf
   └─ Prefill optional values

3. Participant completes questionnaires
   └─ Returns completed PDFs

4. Edit field mapping (optional)
   ├─ "Demographics_Age" → "Age"
   ├─ "Check1" → "Gender"
   └─ Save to configuration

5. Extract data
   ├─ Source: C:\Studies\MyStudy\Completed\P_001/
   ├─ Masterfile: C:\Studies\MyStudy\masterfile.xlsx
   └─ Appends row with mapped column names

File Menu

  • Save Configuration — Save current setup with field mappings
  • Load Configuration — Load previous project setup
  • Delete Configuration — Remove saved configuration
  • Save/Edit Field Mapping — Edit field-to-column-name mappings (auto-scans templates)
  • Exit — Close the application

Generation Section

  • Generate Participant Folder — Create folders with questionnaires and optional prefill

Extraction Section

  • Extract Data to Masterfile — Append completed survey data to masterfile

Technical Details

Version: 2.1 (June 2026)

Changes from v2.0:

  • Field mapping auto-generated from form templates
  • Non-participant-specific column names
  • Integrated field mapping with project configuration
  • Removed redundant load/delete field mapping operations
  • Enhanced type hints and code quality

Technology Stack

  • Python: 3.13.7
  • GUI: PyQt5 5.15.11
  • PDF Processing: pypdf 4.0+
  • Excel Integration: xlwings 0.33.20
  • Packaging: PyInstaller 6.14.2

Supported File Formats

  • PDF: Standard AcroForm fields (text, checkboxes, radio buttons)
  • CSV: UTF-8 encoded, comma-delimited
  • Excel: XLS (2003), XLSX (2007+)

Troubleshooting

"No fields found in configured templates"

  • Verify template PDF paths are correct
  • Ensure PDFs have form fields (AcroForm structure)
  • Check file permissions

"Field mapping does not match configured forms"

  • Forms have changed since configuration was created
  • Click "Save/Edit Field Mapping" to regenerate from current templates
  • Review and adjust friendly names, then save

Excel file not updating

  • Ensure masterfile is not open in Excel
  • Verify path is correct and file is accessible
  • Check that sheet named "Data" exists (or first sheet is used)

Checkbox values not appearing

  • Verify checkbox field names follow pattern: Check{N}_{option}
  • Ensure option numbers are sequential (1, 2, 3...)
  • Check that at least one checkbox is checked in the PDF

Documentation

  • USER_MANUAL.pdf — Comprehensive user guide with screenshots and step-by-step instructions
  • README.md — This file, quick reference and overview

Support

For issues, feature requests, or questions:

  1. Check USER_MANUAL.pdf for detailed guidance
  2. Review configuration in config.ini
  3. Verify file paths and permissions
  4. Consult "Troubleshooting" section above

License

This project is licensed under the MIT License — See LICENSE file for details.

Copyright © 2026 Cagatay Özcan Jagiello Gutt (Original Creator and Developer)

The MIT License permits free use, modification, and distribution while preserving this copyright attribution. Your creative work and engineering effort are permanently credited as the original creator of this tool.


Survey Workbench v2.1 | June 2026 | Open-Source PDF Form Extraction Tool

View on GitHub →


continueAgent

Language: Python
Last updated: 2026-06-25

View README

ContinueAutomation Agent

Intelligent token optimization agent for VS Code Continue extension. Automatically selects between Mistral and Claude based on question complexity, simplifies queries, and tracks costs in real-time.

Features

  • Smart Model Selection: Analyzes question complexity and auto-switches between Mistral (simple) and Claude (complex)
  • Query Simplification: Reduces token usage by simplifying questions while preserving meaning
  • Cost Tracking: Real-time token and cost estimation before and after each query
  • Auto-Install: One-command installation that runs as background service on startup
  • Prompt Caching: Caches frequently used prompts for efficiency

Installation from GitHub

git clone https://github.com/yourusername/continueAgent.git
cd continueAgent
python install_service.py

This will:

  • Install dependencies (Flask, PyYAML, requests)
  • Create startup service (auto-starts on login)
  • Start API server on http://localhost:5001

Configure Continue

Add to ~/.continue/config.yaml:

- name: ContinueAutomation
  provider: openai
  model: continue-automation
  apiBase: http://localhost:5001/v1
  apiKey: dummy

Restart VS Code and select "ContinueAutomation" from the model dropdown.

How It Works

  1. Initial Assessment: Analyzes question, decides model, simplifies query, shows estimated cost
  2. Model Call: Sends simplified question to chosen model (Mistral/Claude)
  3. Final Report: Shows actual token usage and cost

License

MIT License - see LICENSE file

View on GitHub →


tVNS_deviceplan

Language: Unknown
Last updated: 2026-06-10

View README

tVNS_deviceplan

View on GitHub →


GitRef

Language: Python
Last updated: 2026-04-15

View README

GitRef

A git-based reference manager — like Zotero, but your library is a plain git repo.

Install

pip install gitref

Standalone binary

Download from the Releases page:

PlatformAssetInstall
Windowsgitref.exeMove to a folder on your PATH, or run: irm https://raw.githubusercontent.com/CGutt-hub/gitref/main/install.ps1 | iex
LinuxGitRef-x86_64.AppImage or gitrefchmod +x and move to /usr/local/bin/, or run: curl -fsSL https://raw.githubusercontent.com/CGutt-hub/gitref/main/install.sh | bash

From source

pip install git+https://github.com/CGutt-hub/gitref.git

Features

  • DOI / arXiv / ISBN lookup – paste an identifier, metadata is fetched automatically
  • PDF download – auto-downloads from arXiv or DOI resolvers
  • BibTeX store – all references in .resources.bib (human-readable, diffable, standard)
  • Zip archive – PDFs stored compactly in .resources.zip for git efficiency
  • File watcher – gitref watch extracts all PDFs, auto-adds dropped files, auto-repacks on close
  • Clickable links – resources.md has direct PDF links when watcher is running
  • Git sync – auto-commit + pull + push to keep your library in sync across machines
  • Terminal UI – interactive browser with search, detail view, tagging
  • Browser extension – one-click save from any paper page (like Zotero Connector)

Quick start

# Initialise a new library
gitref init

# Add a paper by DOI
gitref add "10.1038/s41586-024-07487-w"

# Add by arXiv ID
gitref add "2401.12345"

# Start watcher (extracts PDFs, watches for new/modified files)
gitref watch

# Open interactive TUI
gitref browse

# Sync with remote
gitref sync

Workflow

DAU-friendly PDF access — no need to use gitref open / gitref close:

  1. Run gitref watch — this extracts all PDFs from the archive
  2. Open resources/resources.md and click any 📎 link to read a paper
  3. Drop a PDF into resources/ and it's auto-added to the library
  4. Close your PDF viewer (or just hit Ctrl+C on the watcher) — files are repacked

Browser extension

GitRef includes a Chrome/Edge extension (Manifest V3) that works like the Zotero Connector — click the toolbar icon on any paper page to save it to your library.

Setup

  1. Start the local server:
    gitref serve
  2. Load the extension:
    • Open chrome://extensions (or edge://extensions)
    • Enable Developer mode
    • Click Load unpacked and select the extension/ folder from this repo
  3. Navigate to any paper page and click the GitRef icon — the reference (and PDF if available) is saved automatically.

The extension auto-detects DOIs, arXiv IDs, and citation metadata from the page, just like Zotero Connector.

Library structure

~/GitRef/
├── .git/
├── .github/workflows/      # auto-regenerates resources.md on push
├── README.md
└── resources/
    ├── .resources.bib       # BibTeX metadata (source of truth)
    ├── .resources.zip       # all PDFs, compressed
    └── resources.md         # browsable table with clickable PDF links

Commands

Run gitref --help for full usage. Key commands:

CommandDescription
gitref initCreate a new library (git repo + bib)
gitref add <id>Add by DOI, arXiv ID, ISBN, or URL
gitref search <q>Search title, authors, tags
gitref listPrint all entries
gitref browseInteractive TUI
gitref watchExtract all PDFs, auto-add/repack on changes
gitref open <key>Extract a single PDF for reading
gitref close <key>Repack a PDF into archive
gitref compactPack all loose PDFs into archive
gitref syncGit add + commit + pull + push
gitref serveStart server for browser extension (port 7342)
gitref exportExport library as RIS

Use -l <path> to specify a custom library location (default: ~/GitRef).

View on GitHub →


labourAIVolt

Language: Python
Last updated: 2026-03-18

View README

labourAIVolt: AI & Human Labour Displacement Analysis for Volt

LAV Analysis

This repository hosts a Nextflow + Python analysis pipeline that automatically fetches current labour-market data from the World Bank public API and quantifies AI-driven labour displacement across the six Volt EU countries with the most active chapters — Germany, France, the Netherlands, Belgium, Italy, and Spain.

The aim is to give Volt Europa and its national chapters an evidence base for labour-market and technology policy: which sectors are shedding jobs fastest as AI and automation accelerate, which countries are most exposed, and how does digital readiness moderate that exposure?

The pipeline is architecturally modelled after the EmotiView project and extends the AnalysisToolbox modular Nextflow framework for scalable, reproducible analysis with automatic result synchronisation.

PlatformRoleContents
GitHubTechnical implementationPipeline, scripts, results
World Bank APIData sourceLive labour-market indicators

Research Background

The labour-market impact of AI and automation is one of the defining policy challenges of the 2020s. Early projections (Frey & Osborne, 2013) estimated that up to 47 % of US jobs faced high computerisation risk; subsequent analyses have moderated that figure while broadening it to task-level disruption rather than wholesale job destruction. What is clear is that the pace and sector distribution of displacement vary substantially across countries depending on industrial structure, education levels, and digital infrastructure.

For a pan-European political movement like Volt, the relevant questions are:

  1. Which EU sectors show the clearest employment decline correlated with automation?
  2. Are Volt's home countries converging toward or diverging from each other on displacement pressure?
  3. Does a country's digital readiness (internet penetration, high-tech exports) buffer it against displacement?
  4. Where should Volt's labour policy — reskilling funds, working-time reform, Universal Basic Income pilots — be concentrated first?

This pipeline operationalises those questions with reproducible, automatically updated data.


Core Research Questions & Hypotheses

  1. Sector displacement ordering: Industry and agriculture will show larger negative employment-share trends than services, consistent with higher Frey & Osborne automation risk for routine physical/cognitive tasks.

  2. Cross-country heterogeneity: Countries with larger manufacturing sectors (Germany, Italy) will exhibit higher AI Displacement Pressure Index (ADPI) than service-dominant economies (Netherlands, Belgium).

  3. Digitalization buffer: Countries with higher Digitalization Readiness Scores (internet penetration + high-tech export share) will show lower net Vulnerability Scores, suggesting that digital transformation simultaneously creates displacement and provides adaptive capacity.

  4. Temporal acceleration: Employment-share trends will steepen post-2018 as AI adoption accelerates across all three sectors, visible as a structural break in the time-series.


Analysis Pipeline

The pipeline is built on the AnalysisToolbox — a modular Nextflow framework for scalable, reproducible data processing with automatic result synchronisation. The LAV-specific pipeline in LAV_analysis/ extends this framework to:

  • Discover country datasets and create per-country output directories (L1).
  • Fetch and parse live labour-market time-series from the World Bank API.
  • Perform standardised normalisation and cleaning.
  • Extract displacement signals and automation-risk-weighted scores per sector.
  • Fit time-series trend models to all available indicators.
  • Aggregate across countries into cross-country rankings and Volt policy metrics (L2).

Configuration is managed via LAV_analysis/LAV_parameters.config. See the AnalysisToolbox documentation for framework details.

Pipeline steps

L1 (per country)
┌─────────────────────────────────────────────────────────────────────┐
│  api_reader              Fetch 13 World Bank indicators (2010–2023) │
│       ↓                                                             │
│  normalizing_processor   Pivot long→wide; sort; deduplicate         │
│       ↓              ↘                                              │
│  displacement_analyzer   trend_analyzer                             │
│  sector scores ×         OLS slope + p-value + R²                  │
│  Frey & Osborne risk     per indicator                              │
└─────────────────────────────────────────────────────────────────────┘
       ↓ collect all countries
L2 (cross-country)
┌─────────────────────────────────────────────────────────────────────┐
│  volt_report_analyzer    Displacement ranking · ADPI · DRS          │
│                          Vulnerability Score · policy metrics       │
└─────────────────────────────────────────────────────────────────────┘

Displacement model

Each broad employment sector receives a displacement score:

displacement_score  =  displacement_signal  ×  automation_risk
TermDefinition
displacement_signalNormalised negative employment-share trend: max(0, −slope / mean_level). A sector losing share faster relative to its baseline scores higher.
automation_riskSector-level probability of computerisation from Frey & Osborne (2013): agriculture 0.82, industry 0.79, services 0.63.
displacement_scoreComposite: high score = fast employment decline and high intrinsic automation susceptibility.

Group-level metrics (L2)

MetricDefinition
ADPI (AI Displacement Pressure Index)Mean displacement_score across all sectors for a country.
DRS (Digitalization Readiness Score)Normalised mean of internet-user percentage and high-tech export share (latest year).
Vulnerability ScoreADPI / (DRS + ε) — high ADPI and low digital readiness = most vulnerable.

Data indicators (World Bank API, no key required)

ColumnWorld Bank codeDescription
employment_agriculture_pctSL.AGR.EMPL.ZSEmployment in agriculture (% total)
employment_industry_pctSL.IND.EMPL.ZSEmployment in industry (% total)
employment_services_pctSL.SRV.EMPL.ZSEmployment in services (% total)
unemployment_rateSL.UEM.TOTL.ZSUnemployment (% labour force)
youth_unemployment_rateSL.UEM.1524.ZSYouth unemployment (%)
employment_to_pop_ratioSL.EMP.TOTL.SP.ZSEmployment-to-population ratio
wage_salary_workers_pctSL.EMP.WORK.ZSWage & salaried workers (%)
internet_users_pctIT.NET.USER.ZSInternet users (% population)
gdp_per_capita_usdNY.GDP.PCAP.CDGDP per capita (current USD)
gdp_growth_annual_pctNY.GDP.MKTP.KD.ZGGDP growth (annual %)
high_tech_exports_pct_mfgTX.VAL.TECH.MF.ZSHigh-tech exports (% manufactured exports)
ict_goods_exports_pctTX.VAL.ICTG.ZS.UNICT goods exports (% total goods exports)
labor_force_totalSL.TLF.TOTL.INTotal labour force

Repository Structure

labourAIVolt/
│
├── LAV_analysis/                    Nextflow pipeline (mirrors EV_analysis/)
│   ├── LAV_pipeline.nf              Main workflow orchestration
│   ├── LAV_modules.nf               IOInterface alias declarations
│   └── LAV_parameters.config        All pipeline parameters & script paths
│
├── LAV_data/                        Per-country input configs (mirrors rawData/)
│   ├── LAV_001/  LAV_001_config.json    Germany        (Volt Deutschland)
│   ├── LAV_002/  LAV_002_config.json    France         (Volt France)
│   ├── LAV_003/  LAV_003_config.json    Netherlands    (Volt Nederland)
│   ├── LAV_004/  LAV_004_config.json    Belgium        (Volt Belgium)
│   ├── LAV_005/  LAV_005_config.json    Italy          (Volt Italia)
│   └── LAV_006/  LAV_006_config.json    Spain          (Volt España)
│
├── LAV_results/                     Pipeline outputs (mirrors EV_results/)
│   ├── .bin/                        Shared infrastructure (logs, HTML archive)
│   ├── LAV_l1/                      First-level: per-country results
│   │   ├── LAV_001/
│   │   │   ├── plots/               Parquet output copies for QC
│   │   │   ├── LAV_001_api_raw.parquet
│   │   │   ├── LAV_001_normalized.parquet
│   │   │   ├── LAV_001_displacement.parquet
│   │   │   ├── LAV_001_trends.parquet
│   │   │   └── LAV_001.log.parquet  Live execution log
│   │   └── LAV_002/ … LAV_006/
│   └── LAV_l2/                      Second-level: cross-country group results
│       ├── LAV_volt_report.parquet
│       ├── LAV_displacement_summary.parquet
│       └── LAV_trends_summary.parquet
│
├── Python/                          Analysis scripts (no Nextflow dependency)
│   ├── lav_run.py                   Standalone orchestrator (used by CI)
│   ├── requirements.txt
│   ├── readers/
│   │   └── api_reader.py            Fetches World Bank labour-market data
│   ├── processors/
│   │   └── normalizing_processor.py Long→wide pivot, clean, sort
│   └── analyzers/
│       ├── displacement_analyzer.py AI displacement scores (Frey & Osborne)
│       ├── trend_analyzer.py        OLS time-series trends per indicator
│       └── volt_report_analyzer.py  Cross-country Volt policy synthesis
│
└── .github/workflows/
    └── lav_analysis.yml             GitHub Actions CI (weekly + on push)

Running the Analysis

The workflow in .github/workflows/lav_analysis.yml runs automatically:

TriggerWhen
ScheduledEvery Monday at 06:00 UTC (pulls the latest World Bank data)
On pushAny change to LAV_data/** or Python/** on main
ManualActions tab → LAV Labour-AI-Volt Analysis → Run workflow

Results are:

  1. Uploaded as a downloadable artifact (lav-results-<run-number>) for 90 days.
  2. Committed back to LAV_results/ in the repository so outputs are versioned alongside the code.

No API keys, secrets, or local software are required.


Option B — Standalone Python (local, no Nextflow)

Use this for quick local runs or debugging individual scripts.

# 1. Clone the repository
git clone https://github.com/CGutt-hub/labourAIVolt.git
cd labourAIVolt

# 2. Install Python dependencies
pip install -r Python/requirements.txt

# 3. Run the full pipeline
python Python/lav_run.py

# Optional: override data/output directories
python Python/lav_run.py --data-dir LAV_data --output-dir LAV_results

Results are written to LAV_results/LAV_l1/<id>/ (per country) and LAV_results/LAV_l2/ (group synthesis).


Option C — Full Nextflow Pipeline (local, requires AnalysisToolbox)

Use this for full pipeline tracing, parallel execution, and integration with the AnalysisToolbox interactive HTML archive.

Prerequisites: Java ≥ 11, Nextflow

# 1. Clone both repos as siblings
git clone https://github.com/CGutt-hub/labourAIVolt.git
git clone https://github.com/CGutt-hub/AnalysisToolbox.git

# Your directory should now look like:
#   parent/
#   ├── AnalysisToolbox/
#   └── labourAIVolt/

# 2. Install Python dependencies
cd labourAIVolt
pip install -r Python/requirements.txt

# 3. Adjust python_exe in LAV_parameters.config if needed
#    (default: 'python3')

# 4. Launch the pipeline from the LAV_analysis/ directory
cd LAV_analysis
nextflow run LAV_pipeline.nf -c LAV_parameters.config

The Nextflow pipeline adds on top of the standalone runner:

  • Parallel per-country execution
  • Full Nextflow trace (LAV_results/.bin/pipeline_trace.txt)
  • Interactive HTML result archive (via AnalysisToolbox interactive_plotter)
  • Automatic git commit + push of results after each country completes

Output Files Reference

Per-country (L1) — LAV_results/LAV_l1/LAV_XXX/

FileDescription
LAV_XXX_api_raw.parquetRaw long-format data as returned by the World Bank API. Columns: participant_id, country, iso3, source, indicator, indicator_code, year, value.
LAV_XXX_normalized.parquetWide-format time-series. One row per year, one column per indicator. Ready for analysis scripts.
LAV_XXX_displacement.parquetPer-sector displacement scores. Key columns: sector, employment_mean_pct, trend_slope_pp_per_yr, trend_significant, automation_risk_frey_osborne, displacement_score.
LAV_XXX_trends.parquetOLS trend results for every indicator. Key columns: indicator, trend_slope, trend_p_value, trend_r_squared, trend_significant, total_change_pct.
LAV_XXX.log.parquetLive pipeline execution log (Nextflow mode only).

Group-level (L2) — LAV_results/LAV_l2/

FileDescription
LAV_volt_report.parquetFull combined table (displacement + policy metrics for all countries).
LAV_displacement_summary.parquetCross-country displacement ranking per sector, with EU-wide mean, std, and per-country rank.
LAV_trends_summary.parquetEU-wide mean slope and significance counts for key indicators across all countries.

Adding a New Country

  1. Create a new directory: LAV_data/LAV_007/
  2. Add a config file LAV_data/LAV_007/LAV_007_config.json:
{
  "participant_id": "LAV_007",
  "country": "Portugal",
  "iso3": "PRT",
  "iso2": "PT",
  "year_start": 2010,
  "year_end": 2025,
  "volt_chapter": "Volt Portugal",
  "population_millions": 10.3,
  "eu_member": true,
  "notes": "Optional notes about the country context"
}
  1. Push the file — the GitHub Action will pick it up automatically on the next run.

Project Status

Active development. Data fetching, pipeline, and group analysis are operational. Planned additions: visualisation layer, structural-break detection (2018 AI inflection point), and integration with OECD employment-by-occupation microdata for finer-grained occupational risk scoring.


References

  • Frey, C. B., & Osborne, M. A. (2013). The Future of Employment: How Susceptible Are Jobs to Computerisation? Oxford Martin School Working Paper.
  • World Bank Open Data. https://data.worldbank.org
  • Acemoglu, D., & Restrepo, P. (2020). Robots and Employment: Evidence from Europe. American Economic Review, 110(6), 2188–2220.
  • Autor, D. (2015). Why Are There Still So Many Jobs? Journal of Economic Perspectives, 29(3), 3–30.

Contributors

NameRoleContact
Cagatay Özcan Jagiello GuttPrincipal InvestigatorORCID

View on GitHub →


paperFinder

Language: Python
Last updated: 2026-02-18

View README

View on GitHub →


Development Philosophy

All code is developed with a commitment to open and transparent science. Tools, pipelines, and analysis code are made available to support reproducibility and collaborative advancement of knowledge.