Sources before conclusions

How Milou builds a traceable analysis

The method starts by making the decision measurable. It then matches each measure to an appropriate source, preserves the source context, performs the analysis, and reports what the evidence cannot establish.

The analysis workflow

  1. Define the decisionClarify the decision, candidate options, geography, population, comparison group, time horizon, and practical constraints.
  2. Translate it into measuresIdentify observable signals such as population change, household income, industry employment, wages, establishments, prices, commute patterns, or housing costs.
  3. Select and inspect sourcesPrefer the original publisher, confirm definitions and geographic coverage, note the release vintage, and inspect suppression, revisions, and sampling limitations.
  4. Prepare comparable dataAlign units, time periods, geographic boundaries, and category definitions before calculating rates, shares, differences, trends, or scores.
  5. Analyze and challenge the resultLook for outliers, denominators that are too small, conflicting indicators, sensitivity to weights, and conclusions that change under reasonable alternatives.
  6. Report evidence and limitsExplain the result in plain language, retain links and dataset identifiers, label estimates and assumptions, and identify missing information that could change the decision.

How sources are selected

A relevant dataset is not automatically the right dataset. Source selection considers authority, directness, geographic resolution, topical fit, release frequency, historical consistency, and whether the measure’s definition matches the question. When possible, an analysis should use a first-party government or institutional publisher rather than a secondary summary.

For example, the American Community Survey provides annual demographic, social, economic, and housing estimates. The Quarterly Census of Employment and Wages publishes employment and wage counts by industry and geography. The Local Area Unemployment Statistics program covers labor force, employment, and unemployment for local areas. Each measures something different; substituting one for another without checking the definition can create a confident but invalid comparison.

Source breadth follows the decision

No fixed source list fits every question. Milou can draw on additional catalog-backed public-data families when their measures match the decision, while keeping each source’s coverage and limitations visible.

  • Road trafficFHWA Highway Performance Monitoring System data can provide annual average daily vehicle traffic on covered road segments. Vehicle counts are not pedestrian foot traffic, store visits, or proof of demand.
  • Hazard and environmental screeningFEMA National Risk Index and EPA Facility Registry Service listings can flag modeled hazard exposure and facilities in environmental programs. They do not replace parcel-level flood determinations, title review, or Phase I or Phase II environmental due diligence.
  • Food access and retail contextUSDA ERS Food Environment Atlas measures describe county food access, store availability, and related conditions. Variables have different source years, and descriptive access measures do not establish profitable demand.
  • Health marketsCDC PLACES, HRSA Area Health Resources Files, and the CMS Provider Data Catalog can inform aggregate prevalence, workforce supply, and provider context. Modeled or area-level measures are not individual diagnoses, patient-flow data, or complete measures of care quality.
  • Education and talent pipelinesNCES IPEDS and College Scorecard cover institutions, completions, fields of study, and selected outcomes. Completions are not current applicants or guaranteed local labor supply, and outcome cohorts can lag current conditions.
  • Weather historyNOAA Global Historical Climatology Network Daily provides station observations used for historical summaries and normals. It is not a site-specific forecast, guarantee about extremes, or estimate of operating costs.
  • Federal lease contextGSA federal lease inventory can add public-sector occupancy and rent context. Federal, often fully serviced and negotiated lease terms are not private-market asking rents or exact property comparables.
  • Financial access and lendingFDIC BankFind Suite, the CFPB Consumer Complaint Database, and HMDA data cover institutions and branches, reported complaints, and mortgage applications. Snapshots, consumer allegations, and unadjusted lending outcomes do not by themselves establish financial health, misconduct, or discrimination.

Provenance and reproducibility

Provenance is the chain from a displayed number back to its origin. A useful record includes the publisher, dataset or series name, table or variable identifier where available, observation period, geographic unit, retrieval context, and transformations applied. For derived measures, the calculation should be stated plainly enough that another analyst can reconstruct it.

Example: “Five-year population change” should identify both endpoint estimates, confirm that the geography is comparable in both periods, and show that the percentage was calculated as (latest - earlier) / earlier. It should not be presented as a forecast.

Source links are a path to verification, not a claim that a publisher endorses Milou or the interpretation. Agency pages can also be revised or reorganized, so dataset names and identifiers matter alongside links.

Freshness and release timing

“Latest” is specific to a series. Monthly labor data, annual survey estimates, quarterly price indexes, and multi-year ACS estimates have different release schedules and reference periods. A current analysis can legitimately combine releases with different vintages, but it should label them instead of implying that every observation describes the same date.

  • Observation period: when the measured activity occurred.
  • Release date: when the publisher made the estimate available.
  • Retrieval date: when the analysis accessed it.
  • Revision status: whether values are preliminary, revised, benchmarked, or final.

Geography and comparability

Counties, metropolitan areas, ZIP Code Tabulation Areas, cities, and census tracts are different units. Names and boundaries can change. A metro-level income estimate should not be silently compared with a city-level establishment count, and a ZIP code used for mail delivery is not the same thing as a Census statistical geography.

Spatial analysis can use Census TIGER/Line shapefiles to connect geographic identifiers with boundaries. The Census Bureau notes that the core TIGER/Line files contain geographic codes but not demographic attributes; those attributes must be joined from a compatible dataset.

Uncertainty, models, and rankings

Survey estimates have sampling error. Administrative data can have reporting lags, suppression, classification changes, and incomplete coverage. Composite scores add modeling choices: inputs, weights, normalization, missing-value handling, and outlier treatment can all change a ranking.

For that reason, a score is a screening device, not a verdict. A sound analysis shows the underlying measures and tests whether a result is driven by one extreme input. The Austin sample illustrates two cases: percentage growth from a very small starting population and unusually high density in a low-income institutional tract.

A practical verification checklist

  • Open the primary source and match the dataset, variable, geography, and period.
  • Confirm whether the value is an estimate, count, index, rate, or modeled result.
  • Check units, inflation adjustment, seasonal adjustment, and annualization.
  • Recalculate important derived values and inspect their denominators.
  • Look for agency notices, revisions, suppressed cells, and breaks in series.
  • Test whether a ranking or conclusion survives reasonable alternative assumptions.
  • Pair public data with local facts such as leases, site visits, customer research, and professional advice before committing capital.

Limits of the method

No workflow eliminates uncertainty. Public datasets may not capture current block-level conditions, private competitors, consumer preferences, lease terms, firm-specific exposure, or causal effects. Absence of evidence is not evidence of absence. Milou provides analysis, not a final verdict; its output should be independently reviewed in proportion to the stakes.

See the method applied

Read the Austin market-entry sample, including its source list, calculation choices, and explicit warnings about the tract score.