How results are measured
Every price move is judged the same way: 28 days of sales after it against the 28 days before. These rules are why the numbers on Results can be trusted, and why some rows refuse to show one.
28-day windows
Both windows are 28 whole days: the 28 days before the day of the change, and the 28 days after it. The day of the change itself is left out of both, because it sold at the old price in the morning and the new price after. A push is measurable once the 28th day after it has ended; until then the event reads "day N of 28" and Results counts it under Moves in flight. A move needs at least 3 units sold in the prior 28 days, or "there is nothing to divide by".
Four weeks, because a week has seven days. Thirty days is four weeks plus two, so one window always holds an extra Saturday or an extra Tuesday that the other does not — and a Saturday at a dispensary is not a Tuesday. Four whole weeks either side holds exactly four of each, so whatever a weekday is worth cancels out instead of landing on one side of the comparison.
An event pushed before this changed goes on being measured over the window it was pushed under. Changing the window of a result somebody has already read would rewrite it.
Days left out
A holiday weekend or a 4/20 falls into one window and not the other, and the spike reads as a response to a price that did not move. Those days are dropped from whichever window they fall in.
Always left out: 1 January, 4/20, 4 July, 24 and 25 December, 31 December, and the US holidays whose date moves — Memorial Day, Labor Day, Thanksgiving and the Friday after it, worked out for the year rather than written down. A store adds its own days on Replenishment, under Days left out of sales rates: a day or range every year, or once (a storm closure), shown read-only on Pricing > Setup. Pricing measures by the days that apply to all products; a day set for one brand or category does not change a measurement.
The window is not made longer to replace a dropped day. Extending it would undo the weekday alignment above. Instead both sides are compared per counted day, so a before-window that lost three days and an after-window that lost one still compare like with like — and dropping a day the product sold nothing on changes nothing at all.
Promotions are read the same way: the excluded days come out of the promotion's own days and out of the weeks before it alike. A promotion that happened to cover 4 July is not thereby a better promotion.
Only the lines still at the event's price
An event is measured on the lines still at the price it set. A line a later event moved, or one whose price someone changed in Dutchie directly, leaves the measurement from that point, and the event says how many it lost and to what ("repriced outside the engine on DATE" for a change made in Dutchie). The same goes for a hold-out line: once its price changes, it is no longer compared against. This applies only inside the window. Once the window has closed the event is Measured and holds nothing, and a later change to its products leaves its result as it was.
Clean sales
What a sold line earned is counted the same way in Pricing and in Promo performance, from one shared definition: what the register collected for the line, its list amount less every discount on it. Returned lines and voided orders never count. A $20 item rung up with $2 off earned $18.
A sale counts unless a campaign discount was on it. Loyalty, a manual discount at the register, and a structural discount - one that runs nearly every day with no end date, such as a standing 10% off - are part of the price customers pay, so those sales count at the price actually paid: a $20 item with $2 of loyalty counts as an $18 sale. A sale under a campaign, a discount that starts and stops, is left out, because it measures the campaign rather than the price. Returned sales, voided orders and free items are left out too.
Promo performance decides which discounts are structural, nightly; it tags them Structural on its list. See How performance is measured.
A product that only sold under a campaign, ran out of stock, or was not selling beforehand is left out of the verdict rather than counted against it.
Days on the shelf
Every figure is counted per day the product could actually sell, not per day on the calendar. A product that was in the stockroom for a fortnight did not lose half its demand; it had half as many days to sell on, and a before-and-after that ignores the difference reads a delivery problem as a price result.
Where the days come from: the nightly inventory history, one row per product per day, holding what was sellable that day. Stock in quarantine is not sellable and does not count as being on the shelf.
A day with no record is treated as a day it was on the shelf. The history starts on 30 May 2026, so a window reaching back before that has days nobody recorded. Those days are counted as normal trading rather than refused, and the event says "part of these windows predates the inventory history" beside the figures so the number is not read as a clean one.
The 70% rule
A line off the shelf for more than 30% of the days that can be seen, in either window, is not judged at all. Its units and its gross profit leave the result, and the line says why: "Out of stock 23 of 28 days after the move - not judged."
The same threshold is read everywhere the engine measures:
- On a price event, the line is dropped from the ratios and from gross profit earned, and its Stocked column reads its days in a warning tone.
- In Results, the move is dropped from its bin and counted in the subtitle as "dropped for being out of stock", beside the cost-controlled count.
- On the shelf check, a neighbour that ran out takes no part on either side: a product nobody could buy from did not stand still, it was absent.
- On Promo performance, a promotion with a third or more of its products off the shelf is Can't judge, because a promotion half of whose products were missing is not a promotion that underperformed.
Seventy percent leaves room for the ordinary gap between deliveries. A product that empties over a long weekend still measures.
What it will not do
It does not estimate what would have sold. A rate is enough to compare two windows honestly, and a lost-sales figure would be a guess dressed as evidence.
A rollback is not a move
The event that puts prices back after a rollback is not measured as a raise or a cut. It restores a price; it does not test one, so it never reaches the bins. It shows as Restored, with what it returned in place of a result, and is left out of the Results scorecard too. The same holds for a walk-back draft once it is pushed. The event it undid is measured as usual, or reads "terminated" if the rollback came on or before the 28th day.
A cut rolled back from its day-7 or day-14 warning is the exception: it still counts, as a labelled move read over the whole days it stood rather than 28. The warning fires only on cuts that are losing, so dropping every cut stopped on it would leave the bins holding only the cuts that were allowed to run, and they would learn that cuts pay better than they do. It counts as labelled even when it had a held-back group, because that group was measured over the full window and the cut over fewer days.
Cost held flat
Gross profit is computed on the cost frozen when the event was drafted. Where the product had been invoiced, that frozen cost is the invoice cost, net of any discount on that invoice line - pricing takes its cost from the most recent invoice that names the product. Each line on the event's page shows it under Cost / now, with the invoice or Dutchie field it came from. A move where cost also changed is excluded: a restock reprice is a cost decision, not a price test.
What a sale cost is a different number and is not this one: the measurement reads line_total_cost, what the register recorded at the moment the unit sold. The two answer different questions and neither stands in for the other.
Bins
Measured moves are grouped by size: Deep cut (15% or more off), Moderate cut (5-15%), Small cut (0.5-5%), Small raise (0.5-5%), Moderate raise (5-15%), Deep raise (over 15%). Each bin reports gross profit, units and per-customer change. Holds are not shown. A change under 5% either way counts as unchanged.
Refusing
A bin refuses, showing "not measured", when it has fewer than 12 moves or one product is more than 40% of them. "A row reading not measured is the engine declining to guess, and it is why the numbers beside the other rows can be trusted." The engine allows raises or cuts only from bins that measured.
Controlled, labelled, reconstructed
- Controlled - pushed as an event with a hold-out, so the effect was measured against a control living through the same days. Only these are summed into Gross profit earned.
- Labelled - pushed as an event, so the intent and window are known, but measured against its own prior period. Anything else that changed in those days is in the number too.
- Reconstructed - inferred from a price differing from the day before. It may be a typo, a cost pass-through or a promotion ending.
How the hold-out is drawn. The moving lines are sorted by how much they sold over the last 90 days and paired off - first with second, third with fourth - and one line of each pair is held back at random. Pairing keeps the arms carrying comparable volume, which is the point of sorting at all; drawing at random inside each pair is what stops the same kind of product landing in the same arm every time. Before this, the split took every second line down the volume order, so whatever odd-ranked products in a bucket had in common was in the treatment arm of every experiment on it.
The draw is reproducible: the event records the seed it used, and the same draft drawn again gives the same arms. The event's hold-out control says so - "Paired by volume, drawn at random (seed 3184701)". An event split before this shows nothing there, because it was not drawn this way and should not be read as though it were.
Where a bin holds at least 3 controlled moves, those decide it, for the engine's permission to raise or cut and for every projection and optimized price alike. Each controlled move is judged against its own hold-out: how the priced products changed, divided by how the held-back ones changed over the same days. A cut that gained 10% while its hold-out gained 20% counts as a loss of ground, not a win. The bin's answer is the typical controlled move (a geometric mean), so one very large product cannot decide it. A hold-out that sold nothing, or made no gross profit, in either window cannot be compared against, and that move counts as labelled instead.
Promotions are measured separately, and found two ways. A promotion that knows its discount is a campaign discount with a start and an end (or a switch-off) 7 to 60 days apart: every product that sold under it, with the discount's id on the sale line, is one promotion, dated at the discount's own start and measured over its own days. The rest are inferred: a dip of 5% or more in realised price lasting 7 to 60 days that returns to list, with no discount named on the sales. Where the two find the same promotion on the same product, the one that knows its discount is kept. Discounts that run every day (structural) are part of the price and are never a promotion here, and a free item is never counted. Bundles and basket deals ("4 for $75") are measured in rows of their own, because what one item realises inside them depends on the rest of the basket. A promotion is measured over its own days, from the first discounted day to the last, counting every sale made during it at what the customer paid, except giveaways. That is compared, a day at a time, with the 28 days before it, counted the same way, so a 10-day promotion is not diluted by 18 days of full-price sales after it. Gross profit uses the cost recorded on each sale, the same way Promotion performance does. A promotion that started as the product's cost changed is left out. That says what a discount sold, not what a permanent price would.
The whole shelf
A hold-out answers "what would these products have done anyway". It does not answer "where did the extra units come from", and the two are different questions. The hold-out is drawn from the event's own moving lines, so it sits in the same bucket as the priced ones - the shelf most exposed to a cut pulling units off its neighbours. A cut on one product that takes units from another in the same bucket reads as a win on the first and, if the second is in the hold-out, as a double win.
The check is the bucket's own total. Every stocked product in the same analysis bucket as the event's lines - the priced ones, the held ones and the untouched ones alike - over the same two windows, counting clean sales the same way everything else here does. Then: the priced lines gained so many units, the bucket gained so many, and the difference came off the neighbours.
The share of the gain the bucket did not get decides the wording. At least half is "Most of the priced lines' gain came from their neighbours"; a fifth to a half is "Part of"; under a fifth is "The bucket grew with them". Where the priced lines lost units there is nothing to have taken from anywhere and the tile says so instead.
This is a check at the level of a whole bucket in one store. It cannot say which neighbour lost which sale, and it is not a model of what substitutes for what. A move it flags is still counted in the response bins and still counted in Gross profit earned; the count is reported beside the move count on Results and on the Pricing hub so the size of the problem is visible before anything is done about it.
A product the weekly tier analysis has never placed in a bucket has no shelf to read, and an event made only of those carries no tile.
The shelf in the forecast
The same reading is taken for every move the store has measured, not only events, and the demand forecast learns from it how much of a move's gain or loss the rest of its bucket gave up or picked up. Three things differ from the tile. Every product in the bucket that moved inside either window is left out of the rest, because its sales answer its own price. A neighbour is any product in the bucket that sold in either window, not only one stocked today, since a strain sold out since was still on the shelf then. And products moved on the same day are one reading, not one per product.
Fitted on one BudLogix store's measured moves on 7 October 2026: a raise's lost sales mostly went to the rest of its bucket (about 64%, likely between 8% and all of it, from 16 readings); cuts are unclear (about 16%, from 15); and promotions took close to nothing from their neighbours (about -6%, likely -19% to +8%, from 310).
It is back-tested on whole buckets. Each held-out set of moves is scored on what the moved products and the rest of their bucket earned together, once with the shelf assumed to stand still and once with the share taken off. The share is used only while it does better on both the typical miss and the large ones. On those moves it did: a typical miss of $47 a day against $50, over 154 sets, almost all of it from four flower raises.
The range
Every estimate now carries a range, not only a number. "Expected +$1,240 gross profit over 90 days (80% likely between -$310 and +$2,900)" says two things: what the measured moves point at, and how much they could be out by.
It is worked out on the ratio of gross profit, on a log scale — which is the scale a ratio is symmetric on, because a move that halves demand and one that doubles it are the same size of surprise and on a plain scale they are not. The moves in the bin give a middle and a spread; the spread over the number of moves gives the width.
80% likely means what it says. Four times in five the real answer falls inside that range, if the next moves behave like the ones already measured. The engine reports 80% rather than the 95% a paper would, on purpose: a 95% range on a dozen measured moves is so wide that every move would straddle zero, and a range that never says anything is decoration. Eighty is the point at which "probably" is worth acting on.
Thin evidence widens the range. A product family with two measured moves borrows the spread of the whole catalogue rather than reporting a suspiciously tight band off two numbers.
A range that straddles zero is an experiment
Where the range runs from a loss to a gain, the evidence cannot say which will happen — whatever the middle number looks like. The engine will not call that a recommendation. It moves to Worth learning on the Pricing hub, and the price event says so under the Return figure: "The range straddles zero — this is an experiment, not a recommendation."
That is the point of having a range at all. A number nobody acts differently on is a number on a screen.
Did the engine's forecast hold
Every push freezes what the engine expected of it: the gross profit, the 80% range around it, and how many days that figure was about. When the window closes, the event shows both.
The two numbers are about different lengths of time - a projection speaks over 90 days, a measurement over 28 - so the event restates the promise over the days that judged it before comparing them: "We said +$1,240 over 90 days (80%: -$310 to +$2,900), which is +$386 over the 28 days that measured it. It earned +$400 - inside the range."
Inside the range is the whole test. A number that lands inside is the engine having been honest about what it did not know; one that lands outside is a miss worth reading. Each line of the event carries the same pair, Expected beside Earned.
Results keeps the running score under Forecast vs actual: how many closed events had a frozen forecast, what share of them landed inside their range, and the average miss both with and without its direction.
An 80% range claims to hold four times in five, so the share is read against 80%:
- About 80% - the ranges are honest.
- Well under - they are too narrow, and the engine is surer than its evidence. Read the ranges as wider than they print.
- Well over - they are too wide, and a range that rules nothing out says nothing.
- Under eight closed forecasts - a share, not yet a finding.
Nothing is adjusted automatically from this. Widening or narrowing the bands is a decision somebody makes and records.
An event pushed before forecasts were frozen has none, and says so rather than showing a zero.
One product's units
The ranges above are about the typical move of a size, which is what the engine decides on. BudLogix does not show what one product would sell if it were cheaper, because the store's own history cannot say it yet.
Every time the bins are measured, each past cut is predicted from the other cuts like it - list or promotion, the same category family, a product selling at a similar pace - and never from itself. A range for one product would be shown only where those predictions held the real move at least 70% of the time and were no more than 60 points wide ("+10% to +60%"). On one BudLogix store's history to 25 September 2026 the ranges were honest (they held 81% of 334 cuts) but far too wide: the narrowest, fast-selling flower cut at its list price, ran from -76% to +109%. One product's four weeks are mostly noise.
Price history on a product shows what it actually sold at each price instead.
Where you see this
Results, the measurement tiles on a Price event, the scorecard on Pricing.
Last checked against the product on 10 Oct 2026. This is the same article operators read inside BudLogix.