Repository navigation
fix: Bugfix-Sweep – Worker-Lebensdauer, HTML-Export, RAG-Verdrahtung, Extraktion - #1
Merged
Merged
Conversation
… Extraktion - QThread-Worker werden bis finished gehalten (kein "Destroyed while thread is still running"), closeEvent stoppt und wartet auf Worker, laufende Extraktion wird nicht mehr ersetzt (worker_utils.py) - HTML-Export: gesamter Text wird escaped, Titel escaped, nur http/https/mailto-Links; GUI nutzt den gehärteten ReportExporter - RAG: Indexierung nicht mehr synchron im GUI-Thread; Chat bekommt immer den aktuellen DocumentManager (keine projektfremden Chunks); Relevanz statt Distanz als Konfidenz; Re-Index bettet erst ein, dann ersetzt er - Projekte: neue Projekte übernehmen die gespeicherte LLM-Konfiguration, Öffnen per ID, aktuelles Projekt wird vor dem Wechsel gespeichert, eindeutige Projektordner - Dateien aus Menü/Ordner werden extrahiert, Warteschlange verliert keine Dateien mehr; Dateitypen aus einer gemeinsamen Liste (.pptx/.html) - Extraktion: Excel-Nullwerte, RTF-\uN/Codepages, BOM/UTF-16, PPTX- Folienreihenfolge, Ressourcen werden geschlossen - TXT-Export entfernt nur Markdown-Syntax; YAML-Front-Matter gequotet; Pandoc mit Timeout; Chat rendert Plaintext; Ollama-Verfügbarkeit wird erneut geprüft; Translator-Wortgrenzen; Companion-Notizen per Projekt-ID Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014VFa8w68q2cEYMqHQqfiH2
|
Welcome! 👋 Thanks for your first pull request in this repository. A maintainer will review it soon. Please make sure:
Thanks for contributing! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Zusammenfassung
Code-Review mit anschließenden Fixes. Die schwersten Befunde:
QThread: Destroyed while thread is still running. Ein Worker wurde schon im Abschluss-Slot verworfen, eine laufende Extraktion wurde ersetzt, undcloseEventhat nicht auf Worker gewartet.<img onerror>,<script>oderjavascript:-Links aus LLM-Ausgaben.Fixes
src/gui/worker_utils.py(retain_until_finished,stop_workers). Alle Worker laufen über_start_tracked_worker. Abbruch überrequestInterruption, undcloseEventwartet auf alle Worker.ReportExporter.markdown_to_htmlescaped den gesamten Text und erlaubt nurhttp,httpsundmailtoals Link-Schema. Der Titel wird escaped. Die GUI nutzt diese eine Implementierung.C#,#12undfile_namebleiben erhalten). YAML-Front-Matter wird perjson.dumpsgequotet. Pandoc läuft mit Timeout und klaren Fehlermeldungen..pptx/.html/.htmneu dabei,.odt/.odsentfernt).\uN-Fallback wird korrekt übersprungen, Multibyte-Codepages werden richtig dekodiert.workspace.idim Export, abwärtskompatibel).Tests
python -m pytest -qmit der vollständigenrequirements.txt(langchain, chromadb): 147 passed, 3 skipped (vorher 103 passed).tests/test_bugfixes_2026_10.py.npm test→ 60/60 (2 neue Tests).tests/linux_platform_smoke.py→ ok.Nicht enthalten: eigene Chroma-Collections pro Projekt. Das ist unnötig, weil Dokument-IDs uuid4 sind und jede Abfrage nach IDs gefiltert wird. Ebenfalls nicht enthalten: ein persistenter Extraktions-Cache.
🤖 Generated with Claude Code
https://claude.ai/code/session_014VFa8w68q2cEYMqHQqfiH2
Generated by Claude Code