Supporting Information for “Anthropogenic Heat in Urban Climate Systems: Forcing, Sensitivity, and Feedback”
- Department of Earth and Environment, Boston University, Boston, USA
- Department of Mechanical Engineering, Boston University, Boston, USA
- Department of Transdisciplinary Science and Engineering, Institute of Science Tokyo, Tokyo, Japan
- Department of Risk and Disaster Reduction, University College London, London, UK
- Center for Climate Change Adaptation, National Institute for Environmental Studies, Tsukuba, Japan
- Department of Civil and Environmental Engineering, Hong Kong University of Science and Technology, Hong Kong, China
- School of Geographical Sciences and Urban Planning, Arizona State University, Phoenix, USA
Contents of this file
Text S1. LLM-assisted corpus assembly and information extraction
Text S2. Lessons from the AI-assisted workflow
Figure S1. Flow of records through the workflow, from search to verified values
Table S1. Branches of the finalized search query with per-branch record counts
Table S2. Study-level sensitivity lookup table
Introduction
This review builds upon a compiled corpus of 577 papers. The primary purpose of this Supporting Information (SI) is to document how this corpus was compiled (Text S1; Figure S1; Table S1) and to summarize the insights gained (Text S2). The SI is intended to provide complete transparency so that the corpus can be retraced, independently examined, and, where appropriate, reproduced.

The corpus was assembled with the help of a large language model (LLM). The primary motivation was scale: an LLM can assist in screening the thousands of records returned by Web of Science searches, identifying those relevant to the scope of this review, and then extracting structured information from the retained papers. Throughout the project we used a single LLM family, Claude Opus by Anthropic, adopting the most capable versions available during the project period (4.7 and 4.8). The newer version was adopted as it became available part-way through the project, and we did not observe systematic differences between the two versions for these tasks. Using the most capable models was important for enabling the workflow, but the rigor of the results was ensured by the surrounding procedures, including screening audits, human adjudication, and the sign-off process discussed later. Through this process, we demonstrate how modern artificial intelligence (AI) tools can assist a small research team in screening, organizing, and synthesizing a large body of literature, while also documenting the lessons learned, limitations encountered, and best practices that emerged from this exercise. In this SI, reviewer denotes an author performing an evaluation or adjudication task within the LLM-assisted workflow described in Text S1, rather than a journal peer reviewer.
Beyond documenting the corpus and its compilation, this SI also provides supporting information for the study-level anthropogenic heat sensitivity synthesis presented in the main text. Table S2 lists the studies included in that figure and maps each abbreviated study label to the corresponding study title and DOI. It also reports the number of sensitivity estimates extracted from each study, the study-level geometric mean sensitivity, and the within-study range shown in the figure.
Text S1. LLM-assisted corpus assembly and information extraction
S1.1. Web of Science search strategy
We started by searching the Web of Science Core Collection in two stages: first, a exploratory query was conducted on 21 April 2026, followed by the finalized query on 22 April 2026. The date defines the reproducible search snapshot, because database coverage, indexing, and document metadata can change after a query is completed.
The exploratory query involved multiple steps. It started with a baseline eight-branch query (Table S1), covering modern terminology related to anthropogenic heat and waste heat in atmospheric, weather, and climate contexts. Then, it was expanded with 12 additional branches, representing historical and adjacent terminology identified primarily from pre-2005 source papers. The resulting 20-branch query returned 5,243 records, of which 2,425 were not retrieved by the baseline eight-branch query.
A manual title-and-abstract review of these 2,425 records found that approximately 80% were irrelevant to the scope of this review, including studies of aquatic thermal discharge, engines, and heat-recovery engineering. We therefore decided to retain only 8 of these 12 branches in the finalized query. Of these 8 branches, 6 were kept without modification because they yielded a meaningful outcome, whereas the remaining 2 were narrowed by adding urban or atmospheric context terms.
Thus, the finalized query comprised 16 branches: the 8 baseline branches and 8 additional branches (see Table S1, which also reports the number of records returned by each branch), returning 3,038 records in Web of Science. After de-duplication and minor record cleanup, 3,035 records were retained as input for the next step.
At this stage, we favored breadth over precision, accepting many irrelevant records to reduce the risk of missing relevant ones. We also did not apply document-type filter, so conference proceedings and other non-article records were retained.
S1.2. LLM title–abstract screening
We then screened each record with the help of an LLM according to two main rules. First, a record was retained only if it presented substantive evidence on at least one of the three topics in the forcing–response–feedback framework: forcing (topic I), sensitivity (topic II), or feedback (topic III). Records spanning multiple topics were assigned the corresponding combinations (I+II, II+III, I+III, or all), whereas records mentioning anthropogenic heat flux only incidentally were labeled others, and out-of-scope records were labeled invalid. Second, topic labels were assigned solely on the basis of evidence reported in the paper itself, rather than its framing or citations.
The 3,035 records from the finalized query were screened on 24 April 2026 using Claude Opus (version 4.7), based on the title and abstract. For each record, the LLM returned an inclusion decision (i.e., whether to include or exclude this record), a topic label, a confidence level (high, medium, or low), and a brief explanation of the rationale. The confidence level reflects the LLM’s assessment of how clearly the title and abstract satisfy the rules discussed above: high when the evidence is explicit, medium when it is suggestive but inconclusive, and low when the available information is limited.
The screening process and associated prompts were refined iteratively through testing and comparison with author judgments, particularly to distinguish papers that estimated anthropogenic heat flux from those that merely cited existing estimates and to clarify the treatment of studies that did not examine air temperature effects. The complete prompt history is archived (see Text S1.9) (Li et al. 2026).
Among the 3,035 records, 513 records were recommended by the LLM for inclusion and 2,522 for rejection. Among the 513 recommendations for inclusion, 276 were assigned high confidence, 229 medium confidence, and 8 low confidence; among the 2,522 recommendations for rejection, 2,497 were assigned high or medium confidence, and 25 low confidence.
To compare the LLM results with author judgment, two authors independently labeled a 93-record reference set (\(=44+49\)). The LLM results agreed with the reference set for 74 of the 93 records (79.6%), with similar agreement across the two subsets (35/44 and 39/49). Because the reference set also informed the prompt refinement, this agreement should be interpreted as an internal consistency check rather than an independent estimate of the LLM’s screening accuracy.
S1.3. Human adjudication and reconciliation
The 276 records recommended for inclusion with high confidence were retained without human adjudication at this stage. The reliability of high-confidence screening calls is supported by the recall audit described later in Text S1.5. Human adjudication therefore focused on the 262 less certain cases: 237 (\(=229+8\)) records recommended for inclusion with medium or low confidence and 25 records recommended for rejection with low confidence.
All 262 records were independently reviewed by two reviewers using a web-based interface that displayed the title, abstract, and LLM results, together with an accept/reject decision form and, where relevant, an option to correct the topic label assigned by the LLM. This stage was deliberately not blinded: reviewers could see the LLM output because the objective was not to perform a fully independent human review of all records, but rather to focus expert attention on the relatively small set of records for which the LLM was less confident.
Among the 262 records, the two reviewers reached agreement on 180 records, defined as the same inclusion decision and topic label; 156 of these were retained. The remaining 82 records were reconciled by considering the LLM’s decision, both reviewers’ notes, and an independent reassessment by a third reviewer. In the end, 64 of these 82 records were retained. For each record, the final decision and its basis were documented. The reconciliation was completed and locked on 4 May 2026. Combining these 220 (\(=156+64\)) manually reviewed records with the 276 high-confidence records yielded a collection of 496 papers, which is referred to as the reconciled shortlist.
Note that the topic assigned to each record at this stage is referred to as the reconciled topic label. The extraction step described later (Text S1.6) may identify evidence for different topics within individual parts of a paper, so its results may not always match the reconciled topic label.
S1.4. Citation backfill through backward citation chasing
Older studies often used terminology that differs from the modern language of “anthropogenic heat’’ and are not consistently indexed in current abstract databases. As a result, even a carefully designed keyword search is likely to miss part of this literature. We therefore supplemented the database search by examining the reference lists of papers in the reconciled shortlist and tracing citations back to earlier studies, a process referred to as citation backfill, with particular attention to pre-2000 work.
The citation backfill was conducted manually. For each candidate reference, the authors reviewed the source, extracted its bibliographic information, standardized the metadata, checked for duplicates in the corpus, screened with the same rules, and retained it after applying the same adjudication and reconciliation procedure discussed in Text S1.3. This process added 48 candidate references, of which 1 was later removed as a duplicate, resulting in a backfill-expanded shortlist of 543 papers.
S1.5. Re-screening and recall audit of rejected records
During human adjudication and reconciliation (Text S1.3), we reviewed all records recommended for inclusion with medium or low confidence, as well as all low-confidence rejections. The medium- and high-confidence rejections were therefore the only records that had not yet received human review. We re-screened and audited these records as described below.
We first re-screened the 2,497 medium- and high-confidence rejected records with the same LLM procedure and rules described in Text S1.2. This re-screening procedure identified 154 records for further review and rejected 2,343 records. Applying the inclusion criteria reduced the 154 records to 28 candidates, of which 25 were retained after full-text assessment.
We next used a stratified random audit to estimate how many relevant papers might still have been missed. The rejected records were grouped by the model’s reported confidence, with a larger sampling fraction assigned to the more uncertain medium-confidence group. Using a fixed random seed for reproducibility, we sampled 150 of the 2,343 records rejected during re-screening: 90 high-confidence and 60 medium-confidence rejections.
None of the 90 high-confidence rejections was found to be relevant. This corresponds to a Wilson 95% upper confidence bound of approximately 4% for the miss rate in this group. We therefore accepted the high-confidence rejections without further review.
Among the 60 medium-confidence rejections sampled, 9 were flagged for further review. Although none of the 9 records was confirmed as a relevant paper after full-text assessment, this number was large enough to justify reviewing all remaining medium-confidence rejections. Re-screening based on their titles and abstracts identified 26 further candidates; 12 were retained after full-text review.
After this step, the provisional corpus comprised 580 papers: 543 in the backfill-expanded shortlist, 25 recovered through re-screening, and 12 added through the recall audit.
S1.6. Structured information extraction
The forcing–response–feedback framework developed in this review article relies on specific information such as reported \(Q_F\) values, temperature responses, and feedback mechanisms. Manual collection of these specifics would be slow, uneven, and hard to audit. This is where LLM assistance becomes particularly useful. An LLM can read every full text against a predefined, version-controlled schema (that is, a defined list of fields to extract, each with a specified format) and return structured, machine-checkable information. This helps keep the extraction consistent, transparent, and easy to audit. Moreover, because each paper was associated with a PDF, the LLM could access the full article, including figures, tables, and captions, in addition to the main text. This was particularly important because relevant evidence was often distributed across different components of a paper, including the main text, figures, tables, and captions.
The extracted results included separate sections for forcing, sensitivity, and feedback. The forcing section captured reported \(Q_F\) values, units, spatial and temporal scales, estimation methods, and sectors. The sensitivity section recorded the temperature variable, magnitude, method, and spatial scale. The feedback section described the feedback mechanism, sign, whether \(Q_F\) was prescribed or prognostic, and whether the paper diagnosed a restoring feedback. Each section was assigned a status of found, partial, or not in paper. Only sections assigned found were treated as providing strong evidence. A partial status indicated that the paper discussed or anticipated the topic without fully establishing it and was therefore treated as weak evidence. Note again that these extracted results do not replace the reconciled topic label discussed earlier (Text S1.3). The two may differ because they are based on information from different sources (one from title/abstract and one from full text).
The extraction was not completed in a single step. It began with a pilot extraction on 13 May 2026, followed by the first systematic extraction from 14 May 2026 to 15 May 2026. The same procedure was subsequently extended to late additions as they entered the corpus, including papers recovered through re-screening and recall audits (Text S1.5) and four additional papers recovered during later review (discussed below). Every extraction run, together with its prompt version and model, is archived (Text S1.9).
The extraction was subsequently refined together with continued reviewer sign-off (to be discussed in Text S1.7), with different sections refined differently. The forcing section was unchanged, with newly identified information recorded in notes. The feedback section underwent repeated refinement as our understanding of what constituted valid feedback evidence evolved. The first refinement (referred to as loop-taxonomy reread) separated source and restoring feedbacks, the second refinement (referred to as quantified-parameter criterion) required a feedback parameter, and the last refinement (referred to as admissibility gate) classified feedback as found only when a paper directly reported a feedback parameter or provided sufficient information to derive one from the reported quantities. Feedback parameters clearing the admissibility gate were subsequently normalized to canonical units (Text S1.8). The sensitivity section also underwent three refinements with substantial reviewer involvement (Text S1.7): first, to extract individual numerical estimates together with other relevant information; second, to normalize the sensitivity values to canonical units, namely, K (W m\(^{-2}\))\(^{-1}\), while preserving their original values and units; and third, to check reported sensitivities against independent quantities reported in the same paper. These refinements supported the subsequent verification process (Text S1.8) while preserving the original extraction records.
Before concluding this subsection, two additional points require clarification. First, a naming difference remains in the archived materials. The first extracted section is referred to as forcing in this SI and the main text, but it retains the original name magnitude in the extraction schema, archived model outputs, and sign-off records. This was because the schema was created before the review article settled on forcing as the section name. The archived files were not renamed retrospectively in order to preserve the original audit trail. Consequently, variables beginning with magnitude in the released data correspond to the forcing section.
Second, we explain how the final paper count for the structured extraction was obtained. The preceding steps produced a provisional set of 580 papers. Because the structured extraction relied on full text, we removed 3 records for which no usable full text was available, including one with a PDF that contained no extractable evidence. We also removed 4 citation-backfill records for which neither an abstract nor a full text could be obtained through institutional access or interlibrary loan. 4 papers were then added and extracted using the same procedure: 2 papers that entered the corpus once their full texts were finally retrieved through institutional subscription, and 2 papers published while this review was being prepared, which were added manually. These changes yielded the locked evidence corpus of 577 papers for structured information extraction. 3 systematic reviews of anthropogenic heat (Sailor 2011; Lu et al. 2024; Feng et al. 2025) were tracked separately as a companion set and were not included in the corpus.
S1.7. Reviewer sign-off of the extraction
The structured extraction was reviewed by the authors through a web interface that displayed each paper’s full text alongside its extracted information. Papers were assigned to reviewers according to their expertise. Reviewers could endorse, modify, or reject the extracted evidence and add a note. All review decisions were documented to provide a complete audit trail. When a paper was reviewed more than once, the most recent decision replaced the earlier one. In total, the reviewer sign-off recorded 959 section-level decisions, which exceeds the number of papers because each paper received separate decisions for each of its sections.
The reviewer sign-off was not only a quality-control step but also a source of methodological refinement. As reviewers examined extracted information, their understanding of what constituted valid evidence evolved, revealing limitations in the original extraction framework. More broadly, as evidence accumulated and our conceptual framework evolved, the structure and organization of the review article were adjusted, requiring refinements to the extraction.
These refinements were implemented differently across different sections of the extraction. The forcing section primarily involved collecting reported evidence, and newly identified information could be incorporated in the notes. Therefore, it did not require repeated re-extraction. In contrast, the sensitivity and feedback sections involved quantitative synthesis and therefore required more iterative refinement. For feedback, this refinement was implemented mainly through progressively stricter definitions of what constituted admissible evidence, leading to multiple re-extractions. For sensitivity, the challenge was not only defining admissible evidence but also obtaining reliable numerical estimates, which required substantial human involvement in interpreting texts, digitizing reported values, and applying the evolving criteria. Throughout this process, insights from reviewer sign-off decisions were incorporated into subsequent extraction steps, allowing the extraction criteria to evolve alongside the scientific framework of the review.
S1.8. Verification and normalization
As discussed above, the sensitivity and feedback sections required more iterative refinement than the forcing section because they involved quantitative synthesis and required substantial human interpretation. Therefore, these two sections underwent additional verification.
For sensitivity values, reviewer sign-off resulted in 129 candidates identified through both LLM extraction and human review of the full text. These candidate values then underwent three rounds of verification: LLM-based verification, review by an independent reviewer, and a final reviewer check. After verification, the final study-level synthesis included 81 individual estimates from 44 papers (Table S2), all normalized to a canonical unit of K (W m\(^{-2}\))\(^{-1}\).
The feedback parameters underwent a similar verification and normalization process but at a smaller scale. The final 13 papers that cleared the admissibility gate were normalized to canonical units (W m\(^{-2}\) K\(^{-1}\) for feedback parameters and dimensionless for gain factors; see the main text), and the feedback figure in the main text is built from these records.
S1.9. Reproducibility and data availability
The project was maintained in a version-controlled GitHub repository, while reviewer decisions were collected through a web interface and recorded in the repository. The workflow was designed to be auditable at every step. The counts reported in this SI are not maintained by hand: they are derived from the archived records by an archived script and enforced by automated consistency checks. The accompanying dataset publication (Li et al. 2026) archives the materials needed to reproduce and examine the workflow.
LLM runs were archived with their inputs, outputs, model identifiers, and prompt versions, and each prompt revision was preserved as a separate file. Because the LLM was accessed through a hosted service, repeating the same request may not produce an identical response. Re-running the archived prompts and inputs should reproduce the substance of the results, but not necessarily the exact wording. The archived outputs therefore provide the definitive record of each run.
Text S2. Lessons from the AI-assisted workflow
In our review, we used AI to support a rigorous, multi-author literature review from initial screening through information extraction and final synthesis, rather than merely assisting an individual researcher with coding or exploratory analyses. To our knowledge, reviews assembled in this manner remain rare. Therefore, we conclude this SI by highlighting three key lessons learned and the remaining unresolved issues, with the aim of enabling other teams to evaluate, adapt, and critically assess this workflow.
S2.1. Version control is the essential guardrail
Integrating AI into a scientific workflow introduces a fundamental challenge: AI systems are fast and scalable, but their mistakes can be difficult to detect and downstream analyses might have been built upon them by the time these mistakes are discovered. Similar issues also arise from human errors, although the speed and scale of propagation may differ. Version control provides a safeguard against both: not because it prevents mistakes, but because it limits their impact by preserving a traceable history of changes, as it has long done in software development where Git is now standard practice.
As an example, an early version of the extraction prompt classified papers as containing evidence of anthropogenic heat feedback too broadly, and a manual review of a small subset of papers exposed the problem. The correction was implemented as a new prompt version rather than a modification: affected papers were re-extracted under the revised prompt, each result was linked to the prompt version that generated it, and the superseded run was retained alongside the updated one. The re-analysis even identified three papers requiring reclassification. The original error, the correction, and the evidence supporting the correction were all preserved in the record.
Interestingly, the most consequential silent error in the project originated from a human process rather than from the AI system. A summary statistic derived from the corpus was miscalculated and remained unnoticed for months. However, because every intermediate state had been preserved, the sequence of changes that caused the error could be traced and corrected. We subsequently consolidated all summary statistics into a single authoritative file validated by automatic checks, ensuring that future inconsistencies can be detected as part of the recorded workflow.
These examples highlight that a version-control system is most valuable when it is continuous, capturing not only final outputs but also the intermediate states that lead to them. Scientific publications typically preserve successful analyses while leaving behind discarded approaches, revised decisions, and failed attempts. This selective documentation is partly a consequence of the substantial effort required to maintain complete records, causing important context to disappear simply because it was never recorded. Working with AI helps reduce this documentation burden because effective AI-assisted workflows require explicit prompts, schemas, and model specifications, which are naturally preserved alongside each run.
A version-control system is also most valuable when it is accessible to all contributors. Git provides a powerful framework for tracking changes, but its steep learning curve makes it an impractical interface for many researchers, particularly those without prior experience using it. In a collaborative review involving contributors with different technical backgrounds, requiring everyone to interact directly through Git would create a barrier to participation. Conversely, allowing contributors to work outside the shared record would recreate the very problem that version control is intended to prevent: silent and unrecoverable changes outside the workflow. The challenge is therefore not only to maintain a version-controlled record but also to make that record usable by all participants. To address this challenge, we built a companion web interface that provides access to the same version-controlled record, allowing reviewers to read a paper alongside its extraction, submit decisions, and automatically incorporate those decisions into the shared history.
In summary, this project was designed around version control from the start: one GitHub repository holds everything, including the corpus, search queries, prompts, schema, every AI run, and every review decision. The evidence reported in this manuscript can be traced to the paper from which it originated, the run that produced it, and the reviewer who validated it. The same record keeps human and machine contributions synchronized: the manuscript, schema, and AI prompts evolve together, ensuring that all components operate from the same shared state.
S2.2. Delegate what we can judge well
The record-keeping function described in the previous lesson is an important but largely infrastructural use of AI. The more consequential step is allowing AI to participate directly in research tasks, such as reading papers, extracting information, and evaluating evidence. Such delegation is appropriate only when two conditions are met: (1) the AI is capable of performing the task, and (2) the researchers understand the task well enough to judge both the output and the way it is produced. The second condition carries the real weight: effective oversight is more than inspecting outputs, because a researcher must know in advance what a scientifically valid and useful result looks like, and what a sound path to that result involves.
As an example, screening was a core task in this project, and it satisfied both conditions for AI delegation. Each paper was judged based only on its title and abstract, a task well within the AI’s capability. Indeed, applying a consistent inclusion criterion across thousands of records is precisely where maintaining consistent human judgment becomes challenging, whereas an AI system can apply the same rule consistently at scale. Each AI decision could be evaluated by a reviewer because the inclusion criteria were sufficiently clear and accessible. Although every decision could have been manually verified, doing so would largely negate the efficiency gains of AI-assisted screening. We therefore used targeted auditing to evaluate screening performance. Among 90 sampled high-confidence rejections, none were found to meet the inclusion criteria, placing the residual miss rate below 4% (Text S1.5). This example illustrates that AI delegation is appropriate when researchers possess sufficient understanding of the task to evaluate AI outputs.
In contrast, the task of identifying and evaluating anthropogenic heat feedback evidence initially failed the second condition. At the beginning of the extraction process, we asked the AI to identify papers that studied anthropogenic heat feedback and provide a qualitative description of the reported feedback mechanisms. Although the AI could successfully identify many papers discussing feedback-like effects, we lacked sufficiently clear criteria for determining whether these papers met the scope of this review. After developing the forcing–response–feedback framework, we narrowed the focus to papers that directly reported a feedback parameter or provided sufficient information to derive one from the reported quantities, which established a clear criterion against which AI outputs could be evaluated. This example illustrates that tasks should not be delegated to AI until humans have a clear basis for evaluating the output. Just as a researcher who cannot assess code quality should not blindly rely on AI-generated code, scientific tasks should only be delegated when researchers understand what constitutes a scientifically valid and useful result within the scope of their research.
S2.3. Human in the loop should improve the loop
Once a task is delegated to AI, humans should remain in the loop, not only to provide oversight and assess the quality of outputs, but more importantly, to use their observations to improve the workflow itself. If a person is involved only to inspect each AI output, the cost grows linearly with the corpus size and the workflow gains little from the interaction. If, instead, a person’s observation changes how the AI operates in subsequent iterations, each review decision becomes an opportunity to improve the workflow.
One example was the extraction of sensitivity values from the papers. Early in the extraction process, a reviewer auditing the sensitivity results identified papers that had been assigned sensitivity values despite ambiguous evidence. In some cases, the paper described a temperature response associated with urbanization broadly or with a combination of urban parameters, without isolating the contribution of anthropogenic heat flux. This observation led to the development of criteria defining whether a reported sensitivity value could be considered attributable to anthropogenic heat flux. These criteria included requirements on the reported temperature response, the attribution to anthropogenic heat flux rather than broader urban effects, and consistency between the spatial and temporal scales of the temperature response and anthropogenic heat flux perturbation. The criteria were incorporated directly into the extraction prompt, allowing subsequent extraction runs to apply them systematically. The result was not merely a correction to individual records, but an improvement to the instrument itself.
The same mechanism operated in other parts of the project as well. For example, human review revealed that the original schema did not distinguish whether anthropogenic heat flux was estimated within the paper or simply adopted from cited references. This distinction was subsequently added as a required field for all extractions. The cumulative impact of this iterative human review was substantial in our project, with 627 of the 959 reviewer sign-off decisions (65%) resulting in changes to the extraction. The ability to propagate such improvements across the corpus is a major advantage of AI-assisted workflows. Once incorporated into a prompt or schema, a revised criterion can be applied consistently across thousands of papers, reducing the possibility of inconsistent interpretation as rules evolve. At the same time, systematic AI errors can reveal where the underlying criteria remain insufficient: repeated error patterns indicate deficiencies in the instrument itself rather than isolated mistakes in individual records, allowing targeted improvements to the workflow.
S2.4. Closing remarks
We recognize that this human-in-the-loop workflow has limitations: AI systems remain partially opaque, runs may not be perfectly reproducible, and human oversight remains essential. The transferability of this approach also remains uncertain because the experience reported here represents a single review team and a single research topic. Nevertheless, the underlying principle appears broadly applicable: route papers by confidence, involve reviewers where scientific judgment is required, record the reasoning behind decisions, and allow those decisions to improve the instrument as well as the outputs. In this model, AI does not replace expert judgment, but extends the reach of expert judgment by applying continuously refined criteria across large bodies of literature.
We are cautiously optimistic about this way of working, not only because of the work it enabled, but also because it suggests an emerging research paradigm based on increasingly capable human–AI collaboration. This workflow enabled a small research team to accomplish work at a scale that would previously have been difficult to achieve. The main investment is not the tooling itself, which can often be adopted within days, but the process of learning how to work effectively with AI as its capabilities continue to evolve. The knowledge and experience gained through this process will continue to carry forward.
| Branch | Records | Beyond baseline |
|---|---|---|
| Baseline branches | ||
| “anthropogenic heat” | 1,206 | — |
| “anthropogenic heating” | 64 | — |
| “waste heat” AND “atmosphere” | 399 | — |
| “waste energy” AND “atmosphere” | 31 | — |
| “waste heat” AND “weather” | 247 | — |
| “waste energy” AND “weather” | 21 | — |
| “waste heat” AND “climate” | 931 | — |
| “waste energy” AND “climate” | 98 | — |
| Baseline branches combined, de-duplicated | 2,818 | |
| Additional branches | ||
| “artificial heat” AND (urban OR city OR atmosphere OR climate OR heat island) | 56 | 54 |
| “man-made heat” OR “manmade heat” | 6 | 5 |
| “heat discharge” AND (urban OR city OR building OR anthropogenic OR atmosphere) | 52 | 27 |
| “waste heat” AND (urban OR city OR building) AND (atmosphere OR climate OR heat island OR urban heat island) | 303 | 15 |
| “metabolic heat” AND (urban OR city OR atmosphere) | 33 | 27 |
| (“heat emission” OR “heat emissions”) AND (urban OR city OR atmosphere OR anthropogenic) | 283 | 85 |
| “human-induced heat” | 6 | 5 |
| “anthropogenic moisture” | 8 | 5 |
| Net addition from the additional branches | \(+220\) | |
| Finalized query output (before de-duplication and cleanup) | 3,038 | |
| Study label | Title | DOI | Type | \(n\) | Geometric mean | Range |
| K (W m\(^{-2}\))\(^{-1}\) | ||||||
| Ghadban (2020) | A novel method to improve temperature forecast in data-scarce urban environments with application to the urban heat island in Beirut | https://doi.org/10.1016/j.uclim.2020.100648 | forcing-based | 1 | 0.002 | 0.002 |
| Cao (2017) | Numerical and experimental study of sensitivity factors on heat island of residence community | https://doi.org/10.1007/s10652-017-9544-x | forcing-based | 1 | 0.00233 | 0.00233 |
| Nakajima (2021) | Human behaviour change and its impact on urban climate: restrictions with the G20 Osaka Summit and COVID-19 outbreak | https://doi.org/10.1016/j.uclim.2020.100728 | effective | 2 | 0.00245 | 0.002–0.003 |
| Lin (2008) | Urban heat island effect and its impact on boundary layer development and land-sea circulation over northern Taiwan | https://doi.org/10.1016/j.atmosenv.2008.03.015 | forcing-based | 1 | 0.003 | 0.003 |
| Khanh (2025) | Impact of anthropogenic heat on air temperature: a first-order estimate using dimensional analysis and numerical simulations | https://doi.org/10.1029/2024GL114400 | forcing-based | 2 | 0.005 | 0.001–0.025 |
| McCarthy (2010) | Climate change in cities due to global warming and urban effects | https://doi.org/10.1029/2010GL042845 | forcing-based | 2 | 0.00632 | 0.003–0.01 |
| Mussetti (2022) | Do electric vehicles mitigate urban heat? The case of a tropical city | https://doi.org/10.3389/fenvs.2022.810342 | forcing-based | 2 | 0.00663 | 0.004–0.011 |
| Huang (2021) | The synergistic effect of urban heat and moisture islands in a compact high-rise city | https://doi.org/10.1016/j.buildenv.2021.108274 | effective | 1 | 0.00678 | 0.00678 |
| Chen (2016) | Model analysis of urbanization impacts on boundary layer meteorology under hot weather conditions: a case study of Nanjing, China | https://doi.org/10.1007/s00704-015-1535-6 | forcing-based | 1 | 0.0069 | 0.0069 |
| Xie (2024) | Could residential air-source heat pumps exacerbate outdoor summer overheating and winter overcooling in UK 2050s climate scenarios? | https://doi.org/10.1016/j.scs.2024.105811 | effective | 2 | 0.00748 | 0.004–0.011 |
| Bueno (2012) | A resistance-capacitance network model for the analysis of the interactions between the energy performance of buildings and the urban climate | https://doi.org/10.1016/j.buildenv.2012.01.023 | effective | 1 | 0.0075 | 0.005–0.01 |
| Karlicky (2026) | Sensitivity of modeled urban climate to urban canopy parameters over central Europe | https://doi.org/10.1088/2752-5295/ae228c | forcing-based | 2 | 0.00792 | 0.0057–0.011 |
| Wang (2023) | Quantifying the impacts of high-resolution urban information on the urban thermal environment | https://doi.org/10.1029/2022JD038048 | forcing-based | 1 | 0.00813 | 0.00813 |
| Xin (2023) | Study of urban thermal environment and local circulations of Guangdong-Hong Kong-Macao Greater Bay Area using WRF and local climate zones | https://doi.org/10.1029/2022JD038210 | effective | 2 | 0.00986 | 0.0021–0.024 |
| Yamaguchi (2025) | Urban cooling and CO2 reduction potentials of mass deployment of heat pump water heaters in Tokyo | https://doi.org/10.1016/j.uclim.2025.102374 | effective | 2 | 0.0109 | 0.007–0.017 |
| Sun (2026) | Modeling urban traffic heat flux in the Community Earth System Model: formulation and validation for two test sites | https://doi.org/10.1029/2025MS005435 | effective | 4 | 0.0128 | 0.008–0.019 |
| Bueno (2011) | Combining a detailed building energy model with a physically based urban canopy model | https://doi.org/10.1007/s10546-011-9620-6 | effective | 2 | 0.014 | 0.013–0.015 |
| Gutman (1975) | Response of the urban boundary layer to heat addition and surface roughness | https://doi.org/10.1007/bf00215641 | forcing-based | 1 | 0.014 | 0.014 |
| Chen (2024) | Modelling the impact of building energy consumption on urban thermal environment: the bias of the inventory approach | https://doi.org/10.1016/j.uclim.2023.101802 | effective | 1 | 0.015 | 0.015 |
| Li XC (2024) | Elevated urban energy risks due to climate-driven biophysical feedbacks | https://doi.org/10.1038/s41558-024-02108-w | effective | 1 | 0.0155 | 0.001–0.03 |
| González-Aparicio (2014) | Impact of city expansion and increased heat fluxes scenarios on the urban boundary layer of Bilbao using Enviro-HIRLAM | https://doi.org/10.1016/j.uclim.2014.07.010 | forcing-based | 2 | 0.0168 | 0.0125–0.0225 |
| Sorbjan (1982) | Some numerical urban boundary-layer studies | https://doi.org/10.1007/BF00124707 | forcing-based | 1 | 0.017 | 0.017 |
| Hu (2012) | Numerical investigation on the urban heat island in an entire city with an urban porous media model | https://doi.org/10.1016/j.atmosenv.2011.09.064 | forcing-based | 1 | 0.017 | 0.017 |
| Fan (2005) | Modeling the impacts of anthropogenic heating on the urban climate of Philadelphia: a comparison of implementations in two PBL schemes | https://doi.org/10.1016/j.atmosenv.2004.09.031 | forcing-based | 8 | 0.018 | 0.002–0.08 |
| Ko (2026) | Modeling the distribution, impacts, and mitigation of anthropogenic heat in Los Angeles | https://doi.org/10.1029/2026JD046326 | forcing-based | 3 | 0.0182 | 0.01–0.03 |
| Yu (1975) | Numerical study of the nocturnal urban boundary layer | – | forcing-based | 1 | 0.0185 | 0.017–0.02 |
| Tao (2022) | Impact of anthropogenic heat emissions on meteorological parameters and air quality in Beijing using a high-resolution model simulation | https://doi.org/10.1007/s11783-021-1478-3 | forcing-based | 4 | 0.0187 | 0.00625–0.0314 |
| Feng (2014) | Impact of anthropogenic heat release on regional climate in three vast urban agglomerations in China | https://doi.org/10.1007/s00376-013-3041-z | forcing-based | 3 | 0.0201 | 0.0169–0.0235 |
| Salvati (2019) | Climatic performance of urban textures: analysis tools for a Mediterranean urban context | https://doi.org/10.1016/j.enbuild.2018.12.024 | forcing-based | 2 | 0.0202 | 0.011–0.037 |
| Myrup (1969) | A numerical model of the urban heat island | – | forcing-based | 1 | 0.023 | 0.023 |
| Lei (2022) | Numerical study of micro-thermal environment in block based on porous media model | https://doi.org/10.3390/buildings12050595 | forcing-based | 1 | 0.025 | 0.025 |
| Neunhäuserer (2007) | Towards urbanisation of the non-hydrostatic numerical weather prediction model Lokalmodell (LM) | https://doi.org/10.1007/s10546-007-9159-8 | forcing-based | 2 | 0.0283 | 0.016–0.05 |
| Li (2013) | A multi-resolution ensemble study of a tropical urban environment and its interactions with the background regional atmosphere | https://doi.org/10.1002/jgrd.50795 | forcing-based | 2 | 0.0287 | 0.006–0.12 |
| Aoyagi (2012) | Numerical simulation of the surface air temperature change caused by increases of urban area, anthropogenic heat, and building aspect ratio in the Kanto-Koshin Area | https://doi.org/10.2151/jmsj.2012-B02 | forcing-based | 2 | 0.0292 | 0.0255–0.0334 |
| Swaid (1991) | Thermal effects of artificial heat sources and shaded ground areas in the urban canopy layer | https://doi.org/10.1016/0378-7788(90)90137-8 | forcing-based | 1 | 0.033 | 0.033 |
| Block (2004) | Impacts of anthropogenic heat on regional climate patterns | https://doi.org/10.1029/2004GL019852 | forcing-based | 2 | 0.0335 | 0.015–0.075 |
| Atwater (1972) | Thermal effects of urbanization and industrialization in the boundary layer: a numerical study | https://doi.org/10.1007/bf02033921 | forcing-based | 4 | 0.0395 | 0.012–0.11 |
| Li D (2024) | Structural uncertainty in the sensitivity of urban temperatures to anthropogenic heat flux | https://doi.org/10.1029/2024MS004431 | forcing-based | 1 | 0.0415 | 0.003–0.08 |
| Wang (2019) | A modified building energy model coupled with urban parameterization for estimating anthropogenic heat in urban areas | https://doi.org/10.1016/j.enbuild.2019.109377 | effective | 1 | 0.045 | 0.04–0.05 |
| Lyu (2024) | Factors influencing the spatial variability of air temperature urban heat island intensity in Chinese cities | https://doi.org/10.1007/s00376-023-3012-y | effective | 1 | 0.058 | 0.058 |
| Swaid (1993) | Urban climate effects of artificial heat sources and ground shadowing by buildings | https://doi.org/10.1002/joc.3370130707 | forcing-based | 1 | 0.064 | 0.064 |
| Kikegawa (2022) | A quantification of classic but unquantified positive feedback effects in the urban-building-energy-climate system | https://doi.org/10.1016/j.apenergy.2021.118227 | effective | 3 | 0.0804 | 0.016–0.369 |
| Ming (2021) | Numerical investigation on the urban heat island effect by using a porous media model | https://doi.org/10.3390/en14154681 | forcing-based | 1 | 0.083 | 0.083 |
| Jacobson (2014) | Effects of biomass burning on climate, accounting for heat and moisture fluxes, black and brown carbon, and cloud absorption effects | https://doi.org/10.1002/2014JD021861 | forcing-based | 1 | 1.26 | 1.24–1.28 |