hg.frequent_itemsets
Frequent-itemset mining, augmented so each transaction can carry a numeric value that is accumulated alongside the item count (support).
This lets you mine itemsets weighted by an arbitrary per-transaction quantity
(revenue, duration, …) in addition to plain co-occurrence support. Pass no
transaction_values to get ordinary (count-only) frequent-itemset mining.
The implementation is Eclat (depth-first search over item tidlists — the sets of transaction indices an item appears in), which is both correct and simple: an itemset’s support is the size of the intersection of its items’ tidlists, and its accumulated value is the sum of those transactions’ values.
(Earlier versions of this module vendored an FP-growth fork whose conditional tree was subtly wrong; this Eclat implementation replaces it and is verified against a brute-force reference.)
- class hg.frequent_itemsets.FrequentItemset(items: list, support: int, value: float)[source]
A frequent itemset result.
Tuple-unpackable as
(items, support, value):items: the items making up the itemset (a list).support: number of transactions containing the itemset (the count).value: the accumulated per-transaction value over those transactions (equalssupportwhen notransaction_valueswere supplied).
- items: list
Alias for field number 0
- support: int
Alias for field number 1
- value: float
Alias for field number 2
- hg.frequent_itemsets.find_frequent_itemsets(transactions: Iterable[Iterable], *, transaction_values: Sequence | None = None, minimum_support: int = 2)[source]
Find frequent itemsets in
transactions, yielded asFrequentItemset((items, support, value)).transactionsis any iterable of iterables of hashable items.transaction_valuesis an optional parallel sequence of numeric weights (one per transaction); when omitted, every transaction has weight 1 sovalueequalssupport(ordinary frequent-itemset mining).minimum_supportis the minimum number of occurrences for an itemset to be reported.>>> transactions = [ ... ['bread', 'milk'], ... ['bread', 'milk', 'eggs'], ... ['milk', 'eggs'], ... ['bread', 'butter'], ... ] >>> for itemset in sorted( ... find_frequent_itemsets(transactions, minimum_support=2), ... key=lambda it: (-it.support, sorted(it.items)), ... ): ... print(sorted(itemset.items), itemset.support) ['bread'] 3 ['milk'] 3 ['bread', 'milk'] 2 ['eggs'] 2 ['eggs', 'milk'] 2
Pass
transaction_valuesto weight each transaction (here by basket price); the third field accumulates that weight over the matching transactions:>>> prices = [4.0, 9.0, 5.0, 7.0] >>> by_value = { ... tuple(sorted(it.items)): it.value ... for it in find_frequent_itemsets( ... transactions, transaction_values=prices, minimum_support=2 ... ) ... } >>> by_value[('bread', 'milk')] # baskets 0 (4.0) and 1 (9.0) 13.0