Preparing for the ASHRAE Building Energy Modeling Professional (BEMP) credential rewards a different habit than most modeling work does: treating every model as a set of decisions someone else will audit. Instead of drilling equations in isolation, study the two-model discipline of Appendix G practice, the paired calibration metrics in ASHRAE Guideline 14, and the diagnostics your simulation tool already reports. Work through scenarios where the intuitive choice is the wrong one, then verify you can explain why. Administrative details such as eligibility, scheduling, and fees are confirmed on ASHRAE's own site rather than here.
Baseline vs. Proposed: Stop Copying Your Design Into Appendix G
Performance and code-compliance modeling depends on keeping the baseline and proposed models structurally separate. The baseline follows prescribed rules, not your design; blending the two silently inflates or deflates the savings you report.
Appendix G of ASHRAE Standard 90.1 is the reference most performance work leans on. Its logic is deliberate: the baseline building uses specified system types, minimum equipment efficiencies, and default operating conditions, while the proposed building reflects the actual design. A common mistake is carrying a design feature into the baseline because it 'seems fair.' If your proposed design includes demand-controlled ventilation, the baseline generally does not inherit it; the baseline gets what the rules assign.
Worked scenario: a modeler simulating a small office assigns the same high-efficiency chiller to both models to 'keep things comparable.' The proposed-vs-baseline savings shrink, and the design team wrongly concludes the chiller upgrade is not cost-effective. The better decision is to set baseline equipment at the prescribed minimum efficiency and let the efficiency difference appear in the results. Why it matters: the entire point of a two-model comparison is isolating the design's contribution, and any borrowed feature corrupts that isolation in either direction.
- Build the baseline from the prescribed system selection and minimum efficiencies, not from your design's equipment list.
- Apply any required system-type rotation rules before adding design details to the baseline.
- Document, input by input, which values came from the design documents and which came from baseline prescriptions.
Calibration Metrics: NMBE Can Hide What CV(RMSE) Reveals
Guideline 14 calibration practice pairs two metrics: normalized mean bias error (NMBE) captures overall over- or under-prediction, while CV(RMSE) captures month-to-month or hour-to-hour shape error. Checking one without the other produces confidently wrong models.
NMBE measures aggregate bias: a value near zero means the model's total matches utility data, but says nothing about whether energy lands in the right months. CV(RMSE) penalizes timing and shape errors: a model that over-predicts summer cooling and under-predicts winter heating can show acceptable NMBE with a large CV(RMSE). Both metrics are normalized, so their acceptability thresholds are expressed as percentages; consult the edition of Guideline 14 your project references for the exact values, since they differ between monthly and hourly data.
Worked scenario: a calibrated office model shows NMBE of about +6% and CV(RMSE) of about 18% against monthly utility data, using illustrative thresholds near ±5% and 15%. The modeler, seeing the bias 'close enough,' declares the model calibrated and proceeds to a retrofit analysis. The better decision is to treat the high CV(RMSE) as a shape problem, inspect the monthly residual pattern, and discover the model flat-lines in shoulder seasons because occupancy schedules ignore school-term cycles. Why it matters: a retrofit estimate built on a model with correct totals but wrong timing can misattribute savings to the wrong end use entirely.
| Metric | What it measures | What a passing value does NOT prove |
|---|---|---|
| NMBE | Aggregate over- or under-prediction relative to measured data | That seasonal or daily distribution matches reality |
| CV(RMSE) | Size of errors period by period, normalized to the mean | That total consumption is unbiased |
| Benchmark EUI comparison | Whether modeled intensity is plausible for the building type | That end-use breakdowns or load shapes are correct |
Zoning Decisions That Skew Every Downstream Result
Thermal zoning determines which loads share a system, which schedules apply where, and how reheat behaves. Coarse zoning averages away the very differences your analysis is meant to expose.
Zone by orientation and by use, not by convenience. Grouping a west-facing perimeter office with a north-facing core zone forces one setpoint and one schedule onto loads that behave differently, smoothing peaks that matter for equipment sizing and savings estimates. In multi-use buildings, separate zones by occupancy pattern first: a conference room intermittently loaded differs from a continuously occupied open office even at the same orientation. Every schedule, setpoint, and internal gain is assigned at the zone or system level, so zoning errors propagate everywhere.
Check yourself with a diagnostic: compare zone-level heating and cooling energy between adjacent zones. If a west perimeter zone and a north core zone show nearly identical profiles year-round, either your zoning or your schedules are wrong. A plausible mistake is merging a storage room into an office zone 'to reduce model size'; the office schedule then loads the storage space, overstating its energy and understating the whole-building savings fraction. The better decision is to accept modest extra zones where use patterns diverge, and record the simplifications you did make.
- Separate perimeter zones by orientation before considering any coarser grouping.
- Split zones where occupancy schedules, setpoints, or internal gains differ, even within one orientation.
- Sanity-check zone energy profiles against each other; near-identical profiles across different orientations signal an input error.
Schedules and Plug Loads: The Inputs That Drive Shape, Not Just Size
Schedules and plug loads control when energy occurs, which drives demand, rate-based savings, and calibration. They are also the least verifiable inputs, so they deserve explicit assumptions and sensitivity checks.
Annual totals hide the shape of consumption, and shape is where schedules matter. Two models can match annual kWh while one predicts a morning startup spike and the other a flat profile; a demand-charge analysis or a controls-retrofit estimate will produce opposite conclusions from those two models. Plug loads deserve equal skepticism: nameplate ratings routinely overstate actual draw, and literature values vary widely by building type. State the source of every plug-load assumption and test it, because no site measurement exists for most early-stage models.
Trace a concrete example: a lighting-retrofit study uses an office schedule with lights off at 6 p.m. The measured data shows significant evening consumption. A mistake here is to 'fix' calibration by inflating plug loads to absorb the evening energy, which then misattributes that energy in the retrofit analysis. The better decision is to adjust the lighting schedule to reflect actual operating hours, document the evidence from the utility interval data, and re-run. Why it matters: the retrofit saves lighting energy only in hours when lighting actually runs, so the schedule correction changes the savings estimate directly.
- Assign each schedule an explicit source: design documents, measured data, or a stated assumption.
- Run a sensitivity case on plug-load density when it is a large share of the load.
- Compare modeled load shape, not just annual totals, against interval utility data when available.
Unmet Hours and Warnings: Reading Model Diagnostics Instead of Ignoring Them
Simulation engines report unmet hours, capacity warnings, and convergence messages for a reason. These diagnostics flag where the model is asking a system for something it cannot deliver, and they change how you interpret results.
Unmet hours indicate that a zone's setpoint could not be maintained, usually because equipment capacity is insufficient or the load input is extreme. For baseline models this can be legitimate: a prescriptive baseline may genuinely be undersized relative to an aggressive proposed design. But unmet hours in the proposed model usually mean an input or sizing error, and results from unmet periods are not comparable across models. Before trusting any savings number, check that both models meet loads throughout the analysis period, or explain and quantify the exceptions.
Treat warnings as a checklist rather than noise. A scenario worth tracing: a chiller shows repeated capacity warnings in the proposed model during design days, yet the annual energy result looks reasonable. The mistake is to accept the result because the annual total 'looks fine.' The better decision is to investigate whether auto-sizing picked a smaller chiller than the design documents specify, correct the capacity, and observe that peak-demand results shift even though annual energy barely moves. Why it matters: demand-charge savings and equipment selection decisions depend on peak behavior that warnings are pointing at.
- Resolve or justify unmet hours before comparing proposed and baseline results.
- Verify auto-sized capacities against design documents when warnings appear.
- Record which warnings you investigated and what you found; an audit trail is part of the deliverable.
Documentation and QA: Making Your Model Auditable
A model another professional can reproduce is the professional standard. Documentation should capture sources for key inputs, every calibration adjustment, and the version history of the model file itself.
Calibration is inherently an adjustment process, and undocumented adjustments are indistinguishable from curve-fitting. Keep a running log that records each change made during calibration, the reason for it, the metric values before and after, and the evidence supporting the change. This applies equally to modeling for code compliance, where the reviewer must verify that each proposed-design feature was modeled as designed. An input convention, such as a consistent way of modeling server rooms or vestibules, should be stated once and applied everywhere in the project.
Add a lightweight QA pass: before issuing results, re-derive a handful of key numbers by hand, check that EUI and end-use breakdowns are plausible for the building type and climate, and confirm that the reported geometry matches the latest architectural set. A realistic failure mode is model-version drift, where results are reported from a file two revisions behind the drawings. The better habit is to freeze a dated model version at each milestone and record which drawing set it reflects. Why it matters: energy models inform capital decisions, and an undated, unreproducible file cannot support those decisions under review.
- Log every calibration adjustment with its rationale and its before/after metric values.
- Freeze and date model versions against specific drawing sets and weather files.
- Verify a sample of inputs by hand before issuing results.
A Rebuild Exercise, Study Sequence, and Readiness Checks
Consolidate your study with a rebuild exercise that forces one-variable comparisons, then follow an adaptable sequence from concepts to scenarios to timed self-assessment. Score yourself against explicit observations, not feelings of familiarity.
Practical exercise: take a simple box model of a single-story office in your simulation tool. Run it once as-is, then make exactly one change at a time: rezone by orientation; tighten the lighting schedule; raise chiller efficiency; disable economizer control. After each run, record total EUI, cooling energy, heating energy, and unmet hours before resetting. Expected observations: rezoning by orientation shifts cooling from the core to west perimeter zones and may create unmet hours if equipment is fixed-capacity; schedule tightening reduces lighting energy but also reduces internal heating in winter, so heating energy can rise. If you do not observe both effects, check whether your zones share schedules they should not.
Adaptable preparation sequence: (1) one week on two-model logic, reading the baseline rules and summarizing them in your own words; (2) one week on calibration math, computing NMBE and CV(RMSE) by hand from a twelve-month example; (3) two weeks on the rebuild exercise above; (4) one week on documentation practice, writing an assumptions log for a past project; (5) ongoing scenario practice using a free BEMP practice question set; (6) a final timed self-assessment. Adjust the durations to your available weeks; the order matters more than the clock.
Readiness self-check rubric, scored 1 to 5 each: you can state which baseline rules governed a recent model and why; you can compute both calibration metrics from utility data without notes; you can explain a case where good NMBE hid bad CV(RMSE); you can list every assumption in a model you built and its source; you can describe what unmet hours changed in your interpretation of results. Treat consistent 4s as a learning milestone indicating you are ready to move to full timed practice, not as a prediction of any exam outcome.
- Rebuild exercise: one variable per run, with EUI, end-use, and unmet-hour observations recorded each time.
- Six-phase sequence: two-model logic, calibration math, rebuild exercise, documentation, scenario practice, timed self-assessment.
- Readiness check: score five specific abilities 1 to 5; consistent 4s indicate readiness for timed practice.
References and further reading
Use these references to explore the concepts and check the latest information from the relevant organizations.
