Algorithm-Classifier-IsolationForest

 view release on metacpan or  search on metacpan

Changes  view on Meta::CPAN

Revision history for Algorithm-Classifier-IsolationForest

0.7.0   2026-07-08/14:30
        - Add explain_samples() / explain_sample_tagged(), reporting which
          features drove a sample's anomaly score, to both the batch class
          and ::Online. Two methods: 'ablation' (the default) substitutes
          each feature in turn with a stored baseline and reports the score
          drop -- fit() now learns per-feature training-data medians for
          this and persists them with the model (::Online uses the medians
          of its retained window instead, tracking drift for free);
          'path' apportions credit over the splits each tree walk crossed
          (local DIFFI; Carletti, Terzi & Susto 2023, see REFERENCES),
          needing nothing beyond the trees, so it also serves models saved
          before baseline support. Ablation is the default because it
          answers the counterfactual for any scored sample; path credit is
          only sharp for samples that were in the training data.
        - Add the `iforest explain` CLI command exposing explain_samples:
          one output line per (row, feature) pair, most responsible
          feature first, with --method path|ablation, -n for top-N
          features per row, and -t to score first and explain only the
          rows clearing the cutoff.
        - Add fit_from_csv(), which trains directly from a CSV file without
          loading it into RAM (for data sets too large to slurp). Streams
          the file in a census pass, a gather pass that keeps only the rows
          the trees sampled (Floyd's algorithm, O(n_trees*sample_size)
          memory), and -- when contamination is set -- a scoring pass whose
          min-heap yields the exact same threshold the in-RAM learner would.
          The census and gather passes only parse the cells they need (row
          count / column width, and the sampled rows), and the contamination
          scoring pass runs through the C backend when use_c is on, so a
          learned threshold over a large file is seconds rather than minutes.
          The first CSV line is skipped automatically when it holds feature
          names (any non-numeric cell, or a match of stored feature_names);
          header => 1 still forces it.
        - fit_from_csv() now defaults to an "index" gather: a fast block-scan
          census records each row's byte offset so the second pass seeks
          straight to the sampled rows instead of re-scanning the whole file,
          cutting a no-contamination fit of a 2M-row file from ~15s to ~3s.
          The offset table costs 8*n bytes and is dropped (falling back to the
          streaming reader) once it would exceed index_max (default 256 MiB);
          pass index => 0 to force streaming. The contamination scoring pass
          defaults to c_scan => 1, letting the C packer coerce cells instead
          of validating each in Perl (identical result on valid data; a
          non-numeric scored cell becomes 0.0 rather than dying). Under
          missing => 'die', a missing cell is now rejected when it lands in a
          sampled training row rather than during a full up-front scan.
        - fit()/from_json(): _pack_tree, which flattens each tree into the
          packed buffers the C scorer walks, now runs in the C backend
          (pack_tree_xs) instead of a recursive Perl closure that built an
          arrayref and six SVs per node and then flattened the lot through
          a map for pack(). It had grown into the largest single phase of
          an axis-mode fit: 100 trees repack in 0.6ms rather than 20.5ms
          (4.2ms rather than 62ms in extended mode), taking a 10k x 8 fit
          from 44ms to 16ms. Same DFS pre-order numbering, same dense-pack
          rule, byte-identical buffers. Skipped on wide-NV perls, where
          c(size) computed in C doubles would differ from _c() in the last
          ulp -- those keep the pure-Perl packer, as _NV_IS_DOUBLE guards
          elsewhere.
        - fit(): under missing => 'die' (the default) the up-front scan for
          undef cells now runs in the C backend when use_c is on, instead of
          a per-cell Perl loop over the whole training set -- 183ms -> 17ms
          on 400k rows x 4 features, where it had been ~68% of fit() time.
          Same row-major order, so the same offending cell is reported. A
          row that is not an arrayref now reads as missing at column 0 on
          both the C and pure-Perl paths (previously a Perl deref error),
          matching how pack_input_xs already treats one.
        - Doc cleanup.
        - Add t/81-sklearn-real-data.t and the four UCI datasets it uses
          (glass, ionosphere, seeds, wdbc; CC BY 4.0, see t/data/README
          for citations). The existing sklearn comparison only ever ran
          against synthetic blobs, and only on machines with Python.
          sklearn's scores are now checked in beside each dataset, so the
          comparison runs everywhere; where Python is present an extra arm
          re-runs sklearn live and catches drift against the checked-in
          reference. Agreement is required to be no worse than our own
          seed-to-seed agreement, rather than against a per-dataset floor:
          an Isolation Forest is a random estimator, so that spread is the
          ceiling, and measuring against it keeps the thresholds from
          encoding how hard a given dataset is to rank. Also checks that C
          and pure-Perl score identically on 30+ correlated columns, and
          that fit_from_csv detects a real header.

0.6.0   2026-07-09/08:15
        - Implement Online Isolation Forest (Filippo Leveni, Guilherme
          Weigert Cassales, Bernhard Pfahringer, Albert Bifet, Giacomo
          Boracchi (2024)) as the new companion class
          Algorithm::Classifier::IsolationForest::Online and releated
          CLI commands.
        - initial prototype support
        - munging via Algorithm::ToNumberMunger

0.5.0   2026-07-04/14:30
        - Wide-NV perls (-Duselongdouble / -Dusequadmath(maybe? does this
          even exist? but possibly some idiot on github... but would fix
          it in this case as well): the pure-Perl tree builder now rounds
          every value it stores (split points, hyperplane coefficients and
          offsets, impute fills) to C double precision at the same points
          the C builder rounds preserving the seed-for-seed bit-identical
          guarantee across backends. Possible breakages for the tests for
          extended mode where -Duselongdouble is in play may exist so tests
          for those systems where it is are skipped for now.



( run in 1.232 second using v1.01-cache-2.11-cpan-9789f410c06 )