> ## Documentation Index
> Fetch the complete documentation index at: https://docs.dema.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Analyze a suggested experiment

> Learn how to interpret the suggested experiment table in Dema and what each metric represents.

<Info>
  When Dema generates a suggested experiment, it provides a detailed table to
  help you understand how the test is structured and what to expect. This table
  breaks down the test's estimated performance metrics and key characteristics
  of the treatment and control groups.
</Info>

## How Dema selects the markets

Before showing you a suggestion, Dema searches for the treatment/control split that
will give the most reliable read. For each candidate split it runs a series of
**back-tests** on your historical data:

* It builds the synthetic control on past data and checks how well control tracks
  treatment **out of sample** — on periods the model didn't train on.
* It injects a known, artificial lift and checks whether the design **detects it**
  (power) and recovers it **accurately**.
* It runs an **A/A check** — looking for lift when none was applied — to confirm
  the split doesn't flag effects that aren't real (false positives).

Each back-test shifts one period further back through history (a "lookback
window" — a week by default for market selection), and the candidates are ranked
on this combined evidence. The result is the suggested split, together with the
statistics and model-quality diagnostics below so you can judge it yourself.

## Understanding the suggested experiment table

When Dema generates a suggested experiment, it provides a detailed table to help you understand how the test is structured and what to expect. This table breaks down the test’s estimated performance metrics and key characteristics of the treatment and control groups.

Here’s what each column represents:

### Experiment statistics

1. **Estimated spend**

   * The amount of advertising spend that will be removed (in a decrease spend test) or added (in an increase spend test) in the treatment group.
   * This helps you understand the magnitude of the change being tested.

2. **Spend previous period**

   * The amount of advertising spend during the same time period as the planned test, but before the test begins.
   * This serves as a baseline to compare current spending and measure changes.

3. **Estimated total gross sales**

   * The total gross sales expected across all regions in the test, including both treatment and control groups.
   * This provides an overview of total expected revenue during the test period.

4. **Estimated incremental gross sales**

   * The amount of sales expected to be lost (in a decrease spend test) or gained (in an increase spend test) in the treatment group due to the test.
   * This is calculated based on the ROAS or epROAS applied to the estimated spend changes.

5. **Treatment group proportion**

   * The percentage of regions assigned to the treatment group versus the control group.
   * A lower proportion indicates that fewer regions will experience the test intervention, helping to reduce potential disruptions.

6. **Estimated lift due to marketing**

   * The percentage change in performance (e.g., sales, profit) expected as a result of the test.
   * This metric indicates the potential impact of the experiment.

7. **Minimum detectable effect (MDE)**

   * The smallest measurable change that can be reliably detected by the test.
   * This reflects the sensitivity of the test design; smaller MDE values indicate higher sensitivity.

8. **Average detected lift** (shown in the app as **Minimum potentially detectable lift**)

   * The average lift the synthetic-control model **actually recovered** during back-testing at the smallest effect size that reaches the power threshold — a measure of the design's sensitivity (the smallest lift it could potentially detect), not what the test inferred.
   * This is **not** the same as the MDE: the MDE is the smallest effect the design *can* detect, while average detected lift is what the model *measured* at that effect size when a known lift was injected.

9. **Treatment group correlation**

   * The correlation between treatment and control groups in terms of historical performance.
   * Higher correlation values (e.g., above 90%) indicate that the groups are well-matched, which improves the reliability of the test.

10. **Required scaling** (for scaling tests)

    * A helper metric that shows how large the required scaling should be for the test.
    * This metric helps you understand the magnitude of spend changes needed in the treatment group to achieve meaningful results.

## Model quality diagnostics

A close historical match isn't automatically a *good* match — a model can fit the
past so tightly that it fails to predict the future (overfitting). The **Model
quality** view reports diagnostics that distinguish a genuinely reliable split from
an overfit one. Use them to gauge how much to trust a result before acting on it.

| Diagnostic                         | What it tells you                                                                                                                                                                            | What good looks like                                                                                                                                                                   |
| ---------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **False positive rate**            | How often the model detected a lift when none was applied (an A/A test across historical windows).                                                                                           | At or below the experiment's significance level (alpha). Higher means the split may flag effects that aren't real.                                                                     |
| **Lift estimation error**          | How far the model's detected lift was from a known lift injected during back-testing, at the minimum detectable effect.                                                                      | Close to 0%. Higher means the model estimates lift less accurately.                                                                                                                    |
| **Effective control markets**      | The effective number of control markets the synthetic control actually leans on, shown as `~N of M` available controls.                                                                      | Close to the number of available controls. A low number means the comparison rests on just a few markets and is more fragile.                                                          |
| **Out-of-sample prediction error** | On periods the model did **not** train on, how far its prediction of the sales was from actual, as a share of average sales. Unlike an in-sample fit, this can't be inflated by overfitting. | Low — a tight, honest match (e.g. under \~10%). Higher means the model predicts the treatment group poorly out of sample.                                                              |
| **Generalization score**           | In-sample error divided by out-of-sample error across historical windows, scored from 0 to 1.                                                                                                | Higher is better: **close to 1** the match generalizes well; **near 0** means the model fit the past too closely and may not hold up during the test (overfitting).                    |
| **Out-of-sample stability**        | How consistent the out-of-sample error stays as the back-testing window slides through history (a stability score from 0 to 1).                                                              | Higher is better: **above 0.5 is stable, 0.3–0.5 is moderate, under 0.3 is fragile**. Shown as N/A with fewer than 3 lookback windows.                                                 |
| **Back-testing windows**           | How many times the proposed setup was re-tested on earlier slices of history. Each window shifts the test back by one period (a week by default).                                            | Typical: 3. More windows give a more confident read but make the check stricter (fewer setups pass). Raise it (5–10) for stable markets; lower it (1–2) for seasonal or volatile data. |

<Tip>
  Read these together, not in isolation. A trustworthy split has a **low
  out-of-sample prediction error**, a **generalization score near 1**, **stability
  above \~0.5**, a **false positive rate at or below alpha**, and an **effective
  control count** that isn't far below the number of available control markets.
</Tip>

### Geographical targeting

<CardGroup cols={2}>
  <Card title="Treatment" icon="location-dot">
    Lists the regions (e.g., zip codes or commute zones) included in the
    treatment group, where the test intervention will occur (e.g., spend
    increased or decreased).
  </Card>

  <Card title="Control" icon="map">
    Lists the regions included in the control group, which serves as the
    baseline for comparison.
  </Card>
</CardGroup>

<Tip>
  **Download CSV**: Provides a downloadable file with the full list of treatment
  and control regions for detailed review and implementation.
</Tip>

<Note>
  **Control weights.** Each control region carries a weight — how much it
  contributes to the synthetic control. **Every** control region stays listed even
  when its weight is `0`, so you always see the full pool that was considered.
  Regions with a negligible weight are shown as exactly `0` (they don't
  meaningfully back the control), and the [effective control
  markets](#model-quality-diagnostics) count reflects only the regions that
  actually carry weight. In the displayed map, control regions are colored according to the weights they contribute with.
</Note>

## How to use this information

### Interpreting estimated metrics

* **Understand the impact on spend and sales**: Compare the estimated spend and incremental gross sales to gauge the scale and potential outcomes of the test.
* **Plan for changes in overall performance**: Use the estimated total gross sales to anticipate any shifts in overall revenue during the test period.
* **Assess test sensitivity**\*: Review the MDE and treatment group correlation to ensure the test design is robust and likely to yield actionable insights.\*

### Preparing for the te*st*

<Warning>
  To ensure successful test implementation:

  * **Align with your team**: Share the treatment and control group details with your team to ensure everyone is aligned on the regions affected.
  * **Download the targeting details**: Use the CSV file to streamline campaign adjustments in ad platforms.
</Warning>

By understanding the suggested experiment table, you can effectively plan and execute tests that drive meaningful insights into the incremental impact of your marketing efforts.
