Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions Directory.Packages.props
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,7 @@
<PackageVersion Include="PdfPig" Version="0.1.16" />
<PackageVersion Include="MimeKitLite" Version="4.18.1" />
<PackageVersion Include="ACadSharp" Version="3.8.0" />
<PackageVersion Include="AngleSharp" Version="1.8.2" />
</ItemGroup>

<ItemGroup Label="Tests">
Expand Down
4 changes: 3 additions & 1 deletion THIRD-PARTY-NOTICES.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,14 +19,15 @@ Each component remains under its own license.
| SkiaSharp | MIT | https://github.com/mono/SkiaSharp |
| Svg.Skia (incl. Svg, ShimSkiaSharp, ExCSS; text shaping in `Filee.Engines/Vector/OutlinedTextPlayback.cs` adapted from it) | MIT | https://github.com/wieslawsoltes/Svg.Skia |
| HarfBuzzSharp | MIT | https://github.com/mono/SkiaSharp |
| SharpCompress (zstd decompression for engine downloads) | MIT | https://github.com/adamhathcock/sharpcompress |
| SharpCompress (zstd decompression for engine downloads; RAR, 7z and TAR comics) | MIT | https://github.com/adamhathcock/sharpcompress |
| Unhwp | MIT | https://github.com/iyulab/unhwp |
| Markdig | BSD-2-Clause | https://github.com/xoofx/markdig |
| ExcelNumberFormat | MIT | https://github.com/andersnm/ExcelNumberFormat |
| ExcelDataReader | MIT | https://github.com/ExcelDataReader/ExcelDataReader |
| PdfPig (PDF text, layout and pictures) | Apache-2.0 | https://github.com/UglyToad/PdfPig |
| MimeKitLite (e-mail / MIME parsing) | MIT | https://github.com/jstedfast/MimeKit |
| ACadSharp (DWG/DXF reading and writing; includes CSMath and CSUtilities) | MIT, Copyright (c) Albert Domenech | https://github.com/DomCR/ACadSharp |
| AngleSharp | MIT | https://github.com/AngleSharp/AngleSharp |
| pypandoc-hwpx (Pandoc AST mapping and package layout in `Filee.Engines/Hwp/Hwpx`, incl. the `blank.hwpx` reference document) | MIT, Copyright (c) 2024 pypandoc-hwpx Contributors | https://github.com/msjang/pypandoc-hwpx |
| Microsoft.Extensions.* | MIT | https://github.com/dotnet/runtime |
| cu2qu (fontTools): the cubic-to-quadratic approach followed by `Filee.Engines/Fonts/Cff/CubicToQuadratic.cs` | Apache-2.0, Copyright 2016 Google Inc. | https://github.com/fonttools/fonttools |
Expand Down Expand Up @@ -59,6 +60,7 @@ Downloaded from their official release pages when the user chooses to install th
| Ghostscript (conda-forge build) | AGPL-3.0 | https://www.ghostscript.com, https://github.com/conda-forge/ghostscript-feedstock |
| Microsoft Visual C++ Redistributable (for Ghostscript, conda-forge `vc14_runtime`) | Microsoft Visual C++ Redistributable license | https://github.com/conda-forge/vc-feedstock |
| FFmpeg 9.0.2 (gyan.dev "full_build-shared" Windows build; ffmpeg.exe, ffprobe.exe and their libraries, incl. x264, x265, libvpx, LAME, Opus, Vorbis, Theora, OpenCORE AMR) | GPL-3.0-or-later (the build is configured with `--enable-gpl --enable-version3`); FFmpeg itself LGPL-2.1-or-later | https://ffmpeg.org, build: https://www.gyan.dev/ffmpeg/builds/ (source: https://github.com/GyanD/codexffmpeg) |
| calibre (ebook-convert) | GPL-3.0 | https://calibre-ebook.com |

These programs run as separate processes. Their source code is available from the linked projects.

Expand Down
18 changes: 18 additions & 0 deletions build/fetch-engines.ps1
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,8 @@
ghostscript/ Ghostscript (AGPL-3.0, separate program) from conda-forge, EPS/PS <-> PDF, with the Microsoft
C++ runtime DLLs (vcruntime/) copied next to gswin64c.exe
ffmpeg/ FFmpeg (GPL-3.0 build, separate program), video and audio; bin/ffmpeg.exe + bin/ffprobe.exe
calibre/ calibre (GPL-3.0, separate program), extracted from the official MSI via an administrative
install: ebook-convert for LIT, LRF, PDB, ... and MOBI / AZW3 output

Every download is pinned to a version and verified with SHA-256 (src/Filee.Engines/Infrastructure/engines.json,
shared with the app). Downloads are cached in build/.cache.
Expand Down Expand Up @@ -216,4 +218,20 @@ if ('ghostscript' -in $selected) {
Copy-Item (Join-Path $Destination 'vcruntime\*.dll') $target -Force
}

if ('calibre' -in $selected) {
$msi = Get-Engine 'calibre'
$tmp = Join-Path $cache 'calibre-admin'
Reset-Folder $tmp
Write-Host 'unpack calibre (administrative install, no system changes)'
$proc = Start-Process msiexec.exe -ArgumentList @('/a', "`"$msi`"", '/qn', "TARGETDIR=`"$tmp`"") -Wait -PassThru
if ($proc.ExitCode -ne 0) { throw "msiexec /a failed with exit code $($proc.ExitCode)." }

$convert = Get-ChildItem $tmp -Recurse -Filter 'ebook-convert.exe' | Select-Object -First 1
if (-not $convert) { throw 'ebook-convert.exe not found after extracting the MSI.' }
$target = Join-Path $Destination 'calibre'
Reset-Folder $target
Get-ChildItem $convert.DirectoryName -Force | Move-Item -Destination $target
Remove-Item $tmp -Recurse -Force
}

Write-Host "Engines ready in $Destination"
59 changes: 54 additions & 5 deletions docs/ENGINES.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,17 +15,19 @@ Filee chooses engines automatically. Settings → *Engines* shows their status,
| Spreadsheets (ExcelDataReader for XLS) | built in (library) | **XLSX, XLS, ODS, CSV, TSV → XLSX, ODS, CSV, TSV** (one CSV / TSV per sheet) |
| Office Open XML | built in | **DOCM / DOTX / DOTM ↔ DOCX, XLSM / XLTX ↔ XLSX, PPTM / POTX / PPSX ↔ PPTX** (macros removed for macro-free types) |
| Fonts | built in | **TTF, OTF, WOFF, WOFF2, EOT ↔ each other**; CFF (PostScript) outlines become TrueType for TTF and EOT |
| **HWPX writer** | built in | **DOCX (+ DOCM/DOTX/DOTM), XLSX (+ XLSM/XLTX), XLS, ODS, CSV, TSV, PPTX (+ PPTM/POTX/PPSX), PDF, TXT, Markdown → HWPX** (and with rhwp → PDF and images); HTML, ODT, RTF, reStructuredText, LaTeX → HWPX with Pandoc |
| **HWPX writer** | built in | **DOCX (+ DOCM/DOTX/DOTM), XLSX (+ XLSM/XLTX), XLS, ODS, CSV, TSV, PPTX (+ PPTM/POTX/PPSX), PDF, HTML, EPUB, MOBI/AZW3, FB2, HWPX, TXT, Markdown → HWPX** (and with rhwp → PDF and images); ODT, RTF, reStructuredText, LaTeX → HWPX with Pandoc |
| **DOCX writer** | built in | **Markdown, TXT, XLSX, CSV, PPTX (+ variants), PDF → DOCX** |
| PDF text | built in (PdfPig) | **PDF → TXT** (reading order, also two columns; no OCR) |
| E-mail | built in (MimeKit) | **EML → HTML** (headers + body, inline pictures), **TXT**, **ZIP** (the attachments) |
| **E-books** | built in | **EPUB, MOBI/AZW/AZW3/PRC, FB2, HTMLZ, TXTZ → EPUB, HWPX (→ PDF), TXT, HTML, Markdown, FB2, HTMLZ, TXTZ**; HTML, Markdown, TXT, DOCX → EPUB/FB2/HTMLZ/TXTZ; HTML → TXT; comics CBZ/CBR/CB7/CBT/CBC → PDF, EPUB, CBZ; PDF → CBZ; AZW4 → PDF |
| CAD (ACadSharp + built-in renderer) | built in (library) | **DWG ↔ DXF**; DWG/DXF → **PDF** (vector), **SVG**, PNG, JPG, WEBP, TIFF, BMP, GIF, ICO, AVIF |
| rhwp | bundled with the installer (`engines/rhwp`) | HWP/HWPX → PDF, **HWP → HWPX, HWPX → HWP** |
| Archives (7-Zip) | bundled with the installer (`engines/7zip`, ~2.5 MB) + built-in readers | ZIP, 7Z, RAR, TAR (+ GZ/BZ2/XZ/Z/7Z/LZ), CAB, ISO, DMG, … **→ folder, ZIP, 7Z, TAR, TAR.GZ/BZ2/XZ**; ALZ, EGG, lzip read in-process; "Compress into one archive" for any files |
| LibreOffice + H2Orestart + Java | **optional download** (~420 MB, 1.3 GB on disk) | older and rare formats only: DOC, XLS, PPT, RTF and OpenDocument → PDF and to each other, output as DOCX/ODT/ODS/ODP; HWP/HWPX → DOCX/ODT; read-only formats (see below) → PDF and their kind's editable formats |
| Pandoc | **optional download** (~42 MB, 240 MB on disk) | Markdown ↔ DOCX/ODT/RTF, DOCX/ODT/HTML/RTF → Markdown, HTML ↔ DOCX/ODT, reStructuredText and LaTeX ↔ Markdown/HTML/DOCX/ODT (and → RTF, EPUB, TXT) |
| Ghostscript | **optional download** (~20 MB, 31 MB on disk) | EPS/PS → PDF, PDF → EPS/PS, PostScript-based AI → PDF (images through PDF and PDFium) |
| FFmpeg | **optional download** (~100 MB, 272 MB on disk) | **all video and audio**: video ↔ video, video → animated GIF or a still frame, audio extraction, audio ↔ audio, GIF → MP4/WEBM/MOV |
| Calibre | **optional download** (~216 MB, 660 MB on disk) | rare e-book formats: LIT, LRF, CHM, PDB, PML, RB, SNB, TCR, OEB → EPUB; EPUB → MOBI, AZW3, LIT, LRF, PDB, PML, RB, SNB, TCR |

DOCX, XLSX, XLS, ODS and PPTX → PDF need neither Microsoft Office nor LibreOffice: they are read in-process, written as HWPX
and rendered by rhwp. Anything the built-in readers understand (including PDF) is written as DOCX by the DOCX writer.
Expand All @@ -51,8 +53,9 @@ they can be installed or removed later in Settings → *Engines*. A conversion t
engine folder only when complete. Redirects to mirrors (even plain HTTP) are followed, because the hash decides.
- Engines go to `%LOCALAPPDATA%\Filee\engines`: outside the app folder, so updates keep them, and inside Filee's
install root, so uninstalling removes them.
- LibreOffice comes as an MSI and is unpacked with an administrative install (`msiexec /a`): files only, no
registry entries, no admin rights. Help, gallery and most dictionaries are removed afterwards (~500 MB).
- LibreOffice and calibre come as MSIs and are unpacked with an administrative install (`msiexec /a`): files only,
no registry entries, no admin rights. LibreOffice's help, gallery and most dictionaries are removed afterwards
(~500 MB).
- Ghostscript comes from conda-forge (Artifex publishes only an NSIS installer): a `.conda` package is a zip with a
zstd tarball, of which only `Library/bin` is unpacked (SharpCompress; `fetch-engines.ps1` uses Windows' `tar.exe`).
Its fonts and resources are compiled into `gsdll64.dll`. The Microsoft C++ runtime it was built against
Expand Down Expand Up @@ -117,9 +120,16 @@ Readers turn the source into a small document model (`Hwp/Hwpx/HwpxModel.cs`) an
start numbers), task lists, quotes, tables with spans, links, images, footnotes and math. No Pandoc needed.
- **PDF** is read with PdfPig (see *PDF reader* below); DOCM / DOTX / DOTM, XLSM / XLTX and PPTM / POTX / PPSX
are read like DOCX, XLSX and PPTX.
- **HTML, ODT, RTF, reStructuredText, LaTeX** are parsed by Pandoc into its JSON AST (`PandocAstReader.cs`, following
- **HTML** is parsed with AngleSharp (`HtmlReader.cs`, `HtmlCss.cs`): headings, paragraphs, bold / italic /
underline / strike / sub / sup / code / mark, the basic inline CSS (colour, background, font weight, style and
size, text-align, text-indent, page breaks) and simple `tag` / `.class` rules of style sheets, nested lists with
start numbers and types, tables with spans / header rows / borders, pictures (local files and data: URIs; remote
pictures are skipped), links and anchors, quotes, `pre`, `hr` and figures. Scripts, styles, forms and `nav` are
skipped. The encoding comes from the BOM, the XML declaration or `<meta charset>`, else UTF-8 or the system
code page. No Pandoc needed.
- **ODT, RTF** are parsed by Pandoc into its JSON AST (`PandocAstReader.cs`, following
[pypandoc-hwpx](https://github.com/msjang/pypandoc-hwpx)).
- Markdown and Pandoc keep structure only, so page setup comes from the built-in template (A4).
- Markdown, HTML and Pandoc keep structure only, so page setup comes from the built-in template (A4).

Element order and attribute values follow files saved by 한글 where the schema and 한글 disagree, for example:

Expand Down Expand Up @@ -343,6 +353,45 @@ OLE objects, charts, video, form controls, text art (its text is kept), arcs / p
master pages, memos, character ratio, relative size and offset, paragraph borders, picture cropping and rotation.
Password-protected (encrypted) HWPX and DRM-wrapped files stop with a clear error.

## E-books

The built-in e-book engine (`src/Filee.Engines/Ebooks`) reads books into the same document model as the HWPX
writer, so every e-book also reaches HWPX, PDF and images. The document model is written back out by
`XhtmlWriter` (EPUB chapters, HTMLZ, single-page HTML), `PlainTextWriter`, `MarkdownWriter` and `Fb2Writer`.

- **EPUB 2 / 3**: container.xml → OPF → spine; each chapter is read by the HTML reader and starts a new page, links
between chapters become links to bookmarks, pictures come from the package, title / authors / language / cover
from the metadata.
- **EPUB output** is EPUB 3 with an NCX for older readers: `mimetype` first and stored, a navigation document from
the headings, chapters split at level-1 headings and page breaks, one style sheet, pictures in formats every
reader shows (others become JPEG / PNG), a cover page for covers the content does not show.
- **MOBI / AZW / AZW3 / PRC** (`Ebooks/Mobi`, written from the MobileRead wiki's format description): PalmDB
records, MOBI header and EXTH metadata, PalmDOC (LZ77) and HUFF/CDIC compression, MOBI 6 `filepos` links and
`recindex` pictures, KF8 text rebuilt from the skeleton and fragment indexes with `kindle:pos` / `kindle:embed` /
`kindle:flow` references resolved, and plain PalmDOC (TEXtREAd) books. **AZW4** (Print Replica) gives back its PDF.
- **FB2**: nested sections become headings, poems / epigraphs / citations / tables are kept, note links become
footnotes, pictures come from the base64 binaries; the XML declaration's encoding (often windows-1251) is used.
FB2 output nests sections by heading level.
- **HTMLZ / TXTZ** (calibre's zipped formats) are read and written, with their `metadata.opf`.
- **Comics**: CBZ, CBR, CB7, CBT and CBC (a ZIP of CBZ files) are opened with SharpCompress; pages are sorted
naturally ("page2" before "page10"). → PDF has one page per picture sized like the picture (JPEGs are embedded
as they are), → EPUB is fixed-layout, → CBZ repacks. PDF → CBZ renders the pages with PDFium at the preset DPI.
- Books with DRM (Adobe, Apple, Kindle) are refused with a clear message; Filee does not remove DRM. KFX and
Topaz books are recognised and refused too.

## Calibre

calibre's `ebook-convert` (GPL-3.0) is an optional download for the formats Filee does not read or write itself:
LIT, LRF, CHM, PDB, PML, RB, SNB, TCR and OEB → EPUB, and EPUB → MOBI, AZW3, LIT, LRF, PDB, PML, RB, SNB and TCR.
EPUB is the hub, so for example FB2 → AZW3 is FB2 → EPUB (built in) → AZW3 (calibre). Its edges cost 20, so
built-in routes always win where they exist.

- The pinned `calibre-64bit-<version>.msi` from download.calibre-ebook.com is unpacked like LibreOffice; the folder
with `ebook-convert.exe` becomes `engines/calibre`.
- Each run gets its own configuration, cache and temp folders in the job's work directory
(`CALIBRE_CONFIG_DIRECTORY`, ...), so a calibre the user installed is never read or changed; messages are
English (`CALIBRE_OVERRIDE_LANG`) and progress comes from its "34% ..." lines.

## HWP ↔ HWPX

rhwp converts between the two 한글 formats without Hancom Office (`export-hwpx` and `convert`). Anything → HWP goes
Expand Down
2 changes: 2 additions & 0 deletions src/Filee.App/Assets/i18n/en.json
Original file line number Diff line number Diff line change
Expand Up @@ -322,6 +322,8 @@
"engines.package.ghostscript.description": "EPS, PS and older Illustrator files → PDF, PNG and other images, and PDF or SVG → EPS/PS. SVG and AI files saved with PDF compatibility convert without it.",
"engines.package.ffmpeg.name": "FFmpeg (video and audio)",
"engines.package.ffmpeg.description": "Video ↔ video (MP4, MOV, MKV, WEBM, AVI, WMV and more), video → animated GIF or a still image, sound from videos (→ MP3, M4A, WAV …), audio ↔ audio (MP3, M4A, FLAC, WAV, OGG, OPUS and more) and GIF → MP4, WEBM, MOV. Only needed for video and audio.",
"engines.package.calibre.name": "Calibre (Kindle and rare e-book formats)",
"engines.package.calibre.description": "Saves as MOBI and AZW3 for Kindle, and converts rare e-book formats: LIT, LRF, PDB, PML, RB, SNB and TCR, plus CHM and OEB input. EPUB, MOBI, AZW3, FB2 and comics (CBZ, CBR) are read without it.",
"engines.package.size": "Download {0} · {1} on disk",
"engines.progress_detail": "{0} of {1} · {2}/s · {3}",
"engines.eta.estimating": "estimating time left…",
Expand Down
2 changes: 2 additions & 0 deletions src/Filee.App/Assets/i18n/ko.json
Original file line number Diff line number Diff line change
Expand Up @@ -322,6 +322,8 @@
"engines.package.ghostscript.description": "EPS·PS 파일과 예전 일러스트레이터 파일을 PDF·PNG 등 이미지로, PDF·SVG를 EPS·PS로 변환해요. SVG와 PDF 호환으로 저장한 AI 파일은 없어도 변환돼요.",
"engines.package.ffmpeg.name": "FFmpeg (동영상·오디오)",
"engines.package.ffmpeg.description": "동영상 ↔ 동영상(MP4·MOV·MKV·WEBM·AVI·WMV 등), 동영상 → 움직이는 GIF·정지 이미지, 동영상에서 소리 추출(→ MP3·M4A·WAV 등), 오디오 ↔ 오디오(MP3·M4A·FLAC·WAV·OGG·OPUS 등), GIF → MP4·WEBM·MOV 변환을 추가해요. 동영상과 오디오에만 필요해요.",
"engines.package.calibre.name": "Calibre (킨들과 드문 전자책 형식)",
"engines.package.calibre.description": "킨들용 MOBI·AZW3로 저장하고, 드문 전자책 형식(LIT·LRF·PDB·PML·RB·SNB·TCR, CHM·OEB 읽기)을 변환해요. EPUB·MOBI·AZW3·FB2와 만화책(CBZ·CBR)은 없어도 읽을 수 있어요.",
"engines.package.size": "다운로드 {0} · 설치 후 {1}",
"engines.progress_detail": "{0} / {1} · {2}/s · {3}",
"engines.eta.estimating": "남은 시간 계산 중…",
Expand Down
2 changes: 2 additions & 0 deletions src/Filee.App/Assets/i18n/zh-CN.json
Original file line number Diff line number Diff line change
Expand Up @@ -322,6 +322,8 @@
"engines.package.ghostscript.description": "EPS、PS 和旧版 Illustrator 文件 → PDF、PNG 等图片,以及 PDF 或 SVG → EPS/PS。SVG 和以 PDF 兼容方式保存的 AI 文件无需它即可转换。",
"engines.package.ffmpeg.name": "FFmpeg(视频和音频)",
"engines.package.ffmpeg.description": "支持视频 ↔ 视频(MP4、MOV、MKV、WEBM、AVI、WMV 等)、视频 → GIF 动图或静态图片、从视频提取声音(→ MP3、M4A、WAV 等)、音频 ↔ 音频(MP3、M4A、FLAC、WAV、OGG、OPUS 等)以及 GIF → MP4、WEBM、MOV。仅视频和音频转换需要它。",
"engines.package.calibre.name": "Calibre(Kindle 与少见电子书格式)",
"engines.package.calibre.description": "可保存为 Kindle 使用的 MOBI 和 AZW3,并转换少见的电子书格式:LIT、LRF、PDB、PML、RB、SNB、TCR,以及读取 CHM 和 OEB。EPUB、MOBI、AZW3、FB2 和漫画(CBZ、CBR)无需它即可读取。",
"engines.package.size": "下载 {0} · 占用 {1}",
"engines.progress_detail": "{0} / {1} · {2}/s · {3}",
"engines.eta.estimating": "正在估算剩余时间…",
Expand Down
4 changes: 2 additions & 2 deletions src/Filee.App/ViewModels/EngineSetupViewModel.cs
Original file line number Diff line number Diff line change
Expand Up @@ -21,8 +21,8 @@ public EngineSetupViewModel(EngineDownloadService downloads, ILocalizer loc)
foreach (var package in Packages)
{
// Large downloads are offered, not pre-selected: LibreOffice is only needed for older formats (DOC, XLS,
// PPT, OpenDocument), FFmpeg (~100 MB) only for video and audio. Ghostscript is small but only needed for
// EPS / PostScript.
// PPT, OpenDocument), FFmpeg (~100 MB) only for video and audio, calibre (~230 MB) only for rare e-book
// formats and Kindle output. Ghostscript is small but only needed for EPS / PostScript.
package.Selected = !package.IsInstalled
&& Filee.Engines.Infrastructure.EngineDownloads.DownloadSize(package.Package) < PreselectLimit
&& package.Package.Id != "ghostscript";
Expand Down
3 changes: 2 additions & 1 deletion src/Filee.Engines/Cad/CadConverter.cs
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@
// DXF and DWG; the drawing is rendered with SkiaSharp to PDF (vector), SVG and raster images.

using Filee.Core.Conversion;
using Filee.Engines.Infrastructure;
using Filee.Engines.Magick;
using ImageMagick;
using Microsoft.Extensions.Logging;
Expand Down Expand Up @@ -41,7 +42,7 @@ from target in (string[])["pdf", "svg", .. ImageEncoder.Writable]
select new ConversionEdge(source, target),
];

public EngineStatus GetStatus() => EngineStatus.Available("ACadSharp: DWG R13–2018+, DXF");
public EngineStatus GetStatus() => EngineStatus.Available("ACadSharp: DWG R13–2018+, DXF", $"{EngineVersions.BuiltIn} · {EngineVersions.Library("ACadSharp", typeof(ACadSharp.CadDocument))}");

public Task<IReadOnlyList<string>> ConvertAsync(ConversionStep step, IProgress<double>? progress, CancellationToken cancellationToken) =>
Task.Run<IReadOnlyList<string>>(() =>
Expand Down
Loading
Loading