This repository is intended as a proof-of-concept for using a generic
datacard to facilitate reusable analyses. Central to the example is
a simple Python module, cardio, which automates the creation and
reading from a YAML datacard.
|-- analysis
| |-- MakeValidationHists.C # process eicrecon output to make hists
| `-- MakeValidationPlots.C # make plots from output of *hists.C
|-- cardio.py # cardio implementation
|-- config.yml # snakemake workflow parameters
|-- input.yml # example input datacard
|-- README.md # description, quickstart
|-- scripts
| |- cleanup.sh # remove run/output directories
| `- full_snakemake.smk # Snakefile to run pipeline without cards
|-- Snakefile # snakemake workflow
`-- template.yml # template output datacard
Cardio can be used interactively in REPL:
>>> import cardio
>>> in_card = cardio.load_card("input.yml")
>>> in_card["description"]
'26.07.1 NC DIS sample with Q^2 = 100 - 1000 GeV^2 generated by Pythia8.316'
>>>
>>> out_card = cardio.make_card("out/output.yml", "template.yml")
>>> out_card["location"]
'./out'Or in a Snakefile, as demonstrated in this repo.
The following workflow demonstrates how these cards can be used for reproducibility.
- Perform initial run:
snakemake --cores 1 --config hist_out="out_0" plot_out="out_0/plot"
- Rerun using the 1st output card as a new template:
snakemake --cores 1 --config hist_out="out_1" plot_out="out_1/plot" template="out_0/output.yml"
- Inspect the output from both runs to see that they're the same.