Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 22 additions & 6 deletions docs/generators/epsilon_inference.rst
Original file line number Diff line number Diff line change
Expand Up @@ -39,10 +39,25 @@ CSSR
CSSR :cite:`Shalizi2002` starts from an IID model and grows causal states in three phases:

1. **Initialize** — one state for the empty history.
2. **Homogenize** — extend suffixes; split or assign child histories when next-symbol
distributions differ (G-test, :math:`\chi^2`, or total-variation threshold).
3. **Determinize** — split homogeneous states until transitions are unifilar; drop
transient bottom-SCC states.
2. **Homogenize** — extend each suffix one symbol into the past, up to ``Lmax``; a
child suffix whose next-symbol distribution differs significantly from its
state's (G-test, :math:`\chi^2`, or total-variation threshold) moves to the best
matching state, or starts a new one. States keep suffixes of every length.
3. **Determinize** — drop transient states, then split states until each state and
symbol lead to a single successor, then keep the most-visited recurrent class.

A length-``Lmax`` suffix has no one-symbol extension in the suffix tree, so its
successor drops the oldest symbol. For a non-Markovian process that can forget the
phase: in the even process with ``Lmax = 3``, the successor of ``011`` on ``1`` would
be the ambiguous ``111``. So the length-``Lmax + 1`` suffix (here ``0111``) is tested
against the truncated suffix's state, and is sent to the best matching state when
the two differ.

Choose ``Lmax`` at least the synchronization length of the source (its order, for
a Markov source). Much larger values run many more significance tests, and some
split states by chance; lowering ``alpha`` counters this. A process that is not
exactly synchronizable has no finite-``Lmax`` reconstruction, and CSSR returns
extra states.

.. autofunction:: cssr

Expand All @@ -51,8 +66,9 @@ Subtree merging

Subtree merging :cite:`CrutchfieldYoung1989` clusters histories with statistically
equivalent next-symbol distributions (metric tolerance ``delta``), then determinizes
to a unifilar presentation. With ``delta=0``, morphs are compared up to a small
numerical tolerance for finite-sample estimates.
to a unifilar presentation. With ``delta=0``, two morphs are equivalent unless a
G-test at significance 0.01 tells them apart, a tolerance that scales with the
sample. Transitions follow the same successor rule as CSSR.

.. autofunction:: subtree_merge

Expand Down
7 changes: 7 additions & 0 deletions docs/generators/stack_inference.rst
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,13 @@ Two families are provided:
Histories are counted with a bounded stack depth, so ``max_stack_depth`` caps
the configurations considered during reconstruction.

Stack CSSR runs flat CSSR over ``(suffix, stack)`` configurations. Every observed
stack seeds its own root, and suffixes grow into the past with the stack fixed.
When morphs are compared, all return symbols count as one event: which return
can follow is decided by the stack top through matched call-return pairs, not by
the finite control. Return edges are matched only to calls observed to close
them.

API
===

Expand Down
Loading
Loading