flowchart LR
A([Idea<br/>Generation]) --> B([Literature<br/>Review])
B --> C([Hypothesis<br/>Formation])
C --> D([Study<br/>Design])
D --> K([Grant<br/>Preparation])
K --> E([IRB<br/>Submission])
E --> F([Data<br/>Collection])
F --> G([Analysis &<br/>Coding])
G --> H([Manuscript<br/>Drafting])
H --> I([Peer<br/>Review])
I --> J([Publication &<br/>Dissemination])
style A fill:#90EE90
style B fill:#FFD700
style C fill:#90EE90
style D fill:#90EE90
style K fill:#FFD700
style E fill:#D3D3D3
style F fill:#D3D3D3
style G fill:#FFD700
style H fill:#FFD700
style I fill:#FF6B6B
style J fill:#D3D3D3
7 Research Domain
The academic medical center’s research enterprise has never been easy to sustain. Investigators spend a growing fraction of their time on grant writing and manuscript administration rather than on investigation. The volume of published literature has outpaced any individual’s ability to track it: PubMed held more than 36.5 million records at the close of fiscal year 2023 and has been taking in between 1.5 and 1.7 million new records a year (National Library of Medicine 2024). The peer review system strains under submission pressure, and the reproducibility crisis — well documented across biomedical fields for more than a decade — continues to surface failures that call the research enterprise’s integrity into question.
Generative AI arrived into this strained system and immediately found traction, not because investigators were looking for AI specifically but because they needed relief from exactly the tasks where AI is most useful: reading and synthesizing large volumes of text, drafting standard-form documents, and generating plausible starting points for complex writing tasks. The technology also introduced new failure modes — citation hallucination, data fabrication, peer review confidentiality violations — that are serious enough to warrant careful institutional guidance.
This chapter maps the real capabilities and real risks of generative AI in the research lifecycle, which means it is about text and code. The AI that matters most to a structural biologist or an imaging researcher is a different technology with different governance questions, and this chapter does not cover it. It does not argue that AI will transform research; the transformation, where it is happening, is more prosaic than that. It argues that an academic medical center that does not give its investigators clear guidance and appropriate infrastructure for AI use in 2025–2026 is leaving a productivity gain on the table and creating an integrity risk at the same time.
The productivity half of that claim deserves its evidence up front, because everything this chapter quantifies afterward is a failure rate, and a reader who sees only those numbers would reasonably conclude the chapter argues against investment. In a preregistered experiment, 453 college-educated professionals were assigned occupation-specific writing tasks, grant writing among them; those given ChatGPT finished 40 percent faster and produced work that graders rated 18 percent higher in quality (Noy and Zhang 2023). The effect is not universal. A randomized trial of 16 experienced open-source developers working in repositories they knew well found the reverse, in a report that has not been peer reviewed: they took 19 percent longer with AI assistance, and estimated afterward that it had made them 20 percent faster (Becker et al. 2025). Read together, the two results describe a gain that is real but conditional, largest where the task is formulaic and the user is not already expert in it, and smallest or negative where both of those are false. They also establish that self-report is not a reliable measure of it, which is worth remembering when investigators tell an institution how much time a tool is saving them.
7.1 The Information Crisis in Biomedical Research
Understanding why AI has taken hold in research requires understanding the scale of the problem it is solving. Publication volume growth is not a recent phenomenon, but the rate has accelerated. A systematic reviewer completing a comprehensive literature search in oncology or cardiology today can expect to screen thousands of abstracts for a single review. The manual process — reading title and abstract, applying inclusion/exclusion criteria, extracting data from included papers — consumes months of researcher time for a single systematic review. A well-conducted systematic review commonly takes upwards of two years from first search to publication, and the methods literature has treated that duration as a problem for as long as it has been measuring it.
Grant writing represents a comparable burden. A survey of 113 astronomers and 82 psychologists active in federal grant seeking found that the average proposal consumed 116 hours of principal-investigator time plus another 55 hours from co-investigators, and estimated that sustained funding rates below roughly 20 percent would drive at least half of active researchers out of federally funded research altogether (Hippel and Hippel 2015). That survey is a decade old and its respondents are not biomedical, so the hours should be read as the shape of the burden rather than as an NIH figure. The shape is familiar enough. A substantial fraction of that time goes into writing that is formulaic: specific aims language, human subjects sections, data management plans, facilities descriptions. These are exactly the document types where AI drafting assistance has the clearest productivity case.
The reproducibility crisis adds a third dimension. Estimates of irreproducibility in preclinical biomedical research range widely, but a 2016 survey published in Nature found that more than 70 percent of researchers had tried and failed to reproduce another scientist’s experiments (Baker 2016). That survey measures what researchers believe about reproducibility rather than the reproducibility rate itself, which is a weaker thing but not a trivial one. Respondents pointed to inadequate methods reporting, insufficient sample sizes, and selective publication; inadequate documentation runs through most of the explanations they offered. AI tools that help investigators write more precise and complete methods sections, generate analysis code that is transparent and version-controlled, and produce CONSORT- or PRISMA-compliant reporting checklists address a real problem.
7.2 Literature Discovery and Synthesis
The most mature and immediately useful AI applications in research are in literature discovery and synthesis. A new generation of “semantic search” tools has emerged that goes well beyond PubMed keyword search: they embed papers as dense vectors, retrieve semantically similar content, and allow natural-language queries over large corpora.
Elicit (elicit.com) allows investigators to upload a set of seed papers or pose a research question and receive a literature matrix — a table in which rows are papers and columns are user-specified attributes extracted by AI (sample size, outcome measures, key findings, population). This is meaningfully different from a list of search results: it is a first-pass data extraction that a reviewer can validate and extend. Elicit’s performance on recall is imperfect — it does not replace a comprehensive database search — but it is a useful starting point for scoping reviews and rapid evidence syntheses.
Consensus (consensus.app) is optimized for answering specific empirical questions: “Does X intervention improve Y outcome?” It returns a meter indicating the degree of published consensus, a list of supporting and contradicting papers, and extracted quotations. It is less useful for exploratory synthesis than Elicit but more useful for answering bounded clinical questions.
SciSpace (scispace.com) and Semantic Scholar (semanticscholar.org) offer complementary capabilities — SciSpace for interactive paper reading with AI explanation, Semantic Scholar for citation network analysis and influence mapping.
Across all these tools, the empirical literature on performance is sobering, though it has to be read with attention to which kind of system was actually tested. A 2024 study in the Journal of Medical Internet Research asked large language models to reproduce the reference lists of eleven published systematic reviews of shoulder rotator cuff pathology, drawn from physiotherapy, sports medicine, orthopedic surgery, and anesthesiology: GPT-4 managed 13.4% precision and 13.7% recall against the references the human reviewers had actually included, and 28.6% of the citations it generated were hallucinated outright; for Bard the hallucination rate was 91.4% (Chelli et al. 2024).
Those numbers describe bare chatbots generating citations from memory, which is not what Elicit or Consensus do. A retrieval-grounded tool returns records from an index, so inventing a paper that does not exist is largely foreclosed by the architecture. Its characteristic failures are quieter: the right paper cited for a claim it does not make, a confident synthesis across studies that in fact disagree, and silent recall failure where the relevant paper was never retrieved at all. Retrieval narrows the problem without closing it. The most careful published evaluation of commercial retrieval-grounded research tools, conducted in law rather than medicine, found hallucination rates between 17 and 33 percent in products marketed as hallucination-free (Magesh et al. 2025). The practical consequence for an institution writing verification guidance is that a rule built only around whether the cited paper exists is calibrated to the chatbot failure and will pass the retrieval failure untouched. For retrieval tools the question is whether the paper supports the sentence it has been attached to, and answering it means reading the paper rather than resolving the DOI.
These findings do not argue against using AI tools for literature work; they argue for using them as a first pass that a human reviews, not as a substitute for comprehensive search.
The practical recommendation for an AMC research computing program is to provide investigators with access to two or three of these tools through an institutional account (most offer institutional pricing), include evaluation guidance in researcher training, and set the expectation that AI-assisted literature search supplements but does not replace structured database search (PubMed, Embase, CENTRAL) for systematic reviews.
7.3 Hypothesis Generation and Study Design
The use of AI in hypothesis generation is the most philosophically interesting application and the one most surrounded by hype. The honest account is more modest.
Language models have been used to generate candidate hypotheses in drug repurposing, protein function prediction, and epidemiological research. In each of these domains, the model’s output reflects statistical patterns in the training corpus — it will propose hypotheses that are plausible given the existing literature, which means it is most useful for generating well-grounded starting points and least useful for generating genuinely novel ones. A model trained on the biomedical literature will not reliably propose hypotheses that contradict established findings, even when those findings are wrong. The value is in speed and breadth: an investigator exploring a new area can spend their judgment on evaluating candidate questions rather than on generating them.
For study design, AI tools are useful for generating draft statistical analysis plans, identifying potential confounders from the literature, drafting power calculations for common study designs, and checking draft methods sections against reporting standards (CONSORT for randomized trials, STROBE for observational studies, PRISMA for systematic reviews). Investigators whose research is an AI model rather than merely assisted by one publish under the AI extensions of those standards, principally TRIPOD+AI for clinical prediction models (Collins et al. 2024) and DECIDE-AI for early-stage clinical evaluation (Vasey et al. 2022); Chapter 19 treats them from the deployment side, and they carry the same authority on the publication side. These are tasks that currently require biostatistician or methodologist time; AI cannot replace that expertise, but it can reduce the number of iterations needed before a statistician review.
The buy-versus-build question here is clearly “use existing tools.” No AMC should be building a hypothesis generation system as a production research service, which is a narrower claim than it may sound: informatics departments studying hypothesis generation as a research question are doing something else entirely and should carry on. For the investigator-facing use cases, current frontier general-purpose models together with domain-specific tools such as Elicit are adequate. The institutional responsibility is to ensure investigators have access to these tools through secure, BAA-compliant channels when the research involves human subjects data.
7.4 Code Generation and Reproducible Analysis
For investigators who work with data — which is most investigators in clinical and translational research — AI code generation is among the highest-value applications in this chapter. The ability to describe a data manipulation task in natural language and receive working R or Python code is genuinely productivity-transforming for investigators who are competent researchers but not expert programmers.
The limitations are important. AI-generated code produces what looks like correct code; whether it is correct depends on whether the investigator can evaluate it. A model asked to perform a mixed-effects regression in R will produce syntactically valid code that often produces numerically plausible results, but it may apply the wrong random effects structure, use the wrong reference level, or fail to account for missing data in the way the investigator intended. An investigator who cannot read the code cannot catch these errors.
The institutional implication is that AI code generation in research requires a floor of computational literacy — the ability to read code, understand its structure, and evaluate whether it implements the intended analysis. This is an argument for including AI-assisted programming in the Training & Workforce Development chapter curriculum, not an argument against using AI for code generation.
Reproducibility is a distinct concern. Research code generated by AI should be committed to version control, documented, and included in data and code availability disclosures. Watermarking is sometimes offered as the technical answer here, and it is worth being precise about what it cannot do. The MarkLLM toolkit is a research framework for watermarking model-generated text, and a watermark is applied at generation time by whoever controls the decoder. An institution cannot retroactively watermark output that a commercial provider generated on a researcher’s personal account, and code is an unusually hostile target because reformatting, refactoring, and linting perturb the token choices that carry the signal. The problem is not that adoption is thin; it is that the mechanism does not fit the use case. What does answer the question of which analysis code was AI-assisted is the institutional gateway log described in Project 2. At minimum, investigators should document in their methods sections whether AI tools were used for analysis code generation, consistent with the disclosure norms discussed below.
7.5 Manuscript Drafting, Authorship, and the Policy Landscape
The use of AI in manuscript preparation has generated more policy activity than any other AI application in research, and the policies from journals and funders converge on a small number of principles that are worth knowing precisely.
AI cannot be an author. The International Committee of Medical Journal Editors updated its authorship recommendations in 2023 to state explicitly that AI tools do not meet authorship criteria because they cannot take responsibility for the work, cannot consent to authorship, and cannot be held accountable for errors (International Committee of Medical Journal Editors 2023). The major biomedical publishers adopted equivalent language within weeks of each other, as Table 7.1 records. The responsible corresponding author is accountable for any AI-generated content in the manuscript, including any errors it contains.
AI use must be disclosed. The specificity of disclosure requirements varies by journal. Cell Press requires a dedicated “Declaration of Generative AI and AI-Assisted Technologies in the Writing Process” section placed before the references. Nature requires disclosure in the methods. JAMA requires disclosure of the tool, the extent of use, and confirmation that the author has verified the final content (JAMA Network 2023). An investigator drafting a manuscript should check the target journal’s specific requirements; the safe default is to describe in the methods or a dedicated section which AI tools were used and for which tasks.
AI in peer review is prohibited. The NIH issued Notice NOT-OD-23-149 in June 2023 prohibiting peer reviewers from using AI tools to analyze or critique NIH grant applications, citing confidentiality requirements for the content of applications (National Institutes of Health 2023). The NSF issued parallel guidance for its merit review process (National Science Foundation 2023). Several large publishers, Springer Nature among them, bar uploading manuscripts to AI tools for review purposes, and ICMJE requires a reviewer to obtain the journal’s permission before using AI at all and to disclose it afterwards (International Committee of Medical Journal Editors 2026). The confidentiality rationale is sound: when a reviewer pastes an unpublished manuscript into a commercial AI service, the manuscript has been disclosed to a third party, potentially in violation of the reviewer’s confidentiality agreement. This is true regardless of whether the AI provider claims not to train on the submitted content.
One caveat belongs on this, for an institution making a five-year infrastructure decision. The confidentiality rationale is about disclosure to a third party, which is a property of where the model runs rather than of the practice itself. A reviewer using a model inside the institution’s own BAA-covered tenant, which is exactly what Project 2 builds, has not disclosed anything to anyone. Current policy does not draw that distinction and an institution should follow current policy. But the rationale the policies give applies to only one of the two cases, and publishers are already piloting AI-assisted triage, so the prohibition is better treated as the present state of the rules than as a settled endpoint.
Grant applications now face an originality bar, not just a disclosure norm. Through mid-2025, NIH’s position on AI in application preparation was disclosure-oriented guidance: the intellectual contribution to the science should be the PI’s own. In July 2025 the agency went further. Notice NOT-OD-25-132 states that NIH will not consider applications substantially developed by AI, or containing sections substantially developed by AI, to be the original ideas of the applicant, and the same notice caps each principal investigator at six applications per calendar year (National Institutes of Health 2025b). The operative move is a finding of non-originality rather than a disclosure penalty, which is why it sits in its own column of Table 7.1. The practical boundary is unchanged — AI as an editing and drafting tool for standard-form sections (data management plans, biosketches, resource descriptions) versus AI as a substitute for the investigator’s scientific thinking — but the consequence of crossing it is no longer a disclosure lapse. It is an application that will not be reviewed.
The rule carries an enforcement problem that the notice does not solve, and an AMC is better off thinking about it before one of its investigators is accused rather than after. No reliable method exists for establishing that an application was substantially developed by AI. The available detectors were built to classify text as machine-generated or not, and they fail in a specific and consequential direction: GPT detectors misclassify writing by non-native English speakers as AI-generated at rates high enough to make detector-based enforcement a discriminatory practice (Liang et al. 2023). The investigators most helped by AI drafting assistance are the same ones most exposed to a false accusation, and at most academic medical centers they are a large share of the postdoctoral and graduate research workforce. An institution cannot fix a federal agency’s evidentiary standard. It can decide in advance what it will tell an investigator who is accused and what documentation would support them, which is a question the secure gateway of Project 2 turns out to answer.
Table 7.1 summarizes the policy landscape for major journals and funders.
| Organization | AI authorship | Disclosure required? | Originality standard | AI in peer review | Policy date |
|---|---|---|---|---|---|
| ICMJE (all member journals) | Prohibited | Yes, extent of use | Not applicable | Journal permission required in advance; use must be disclosed; no upload where confidentiality is not assured | May 2023; AI guidance updated Jan 2026 |
| Nature | Prohibited | Yes, in methods | Not applicable | Prohibited | Jan 2023 |
| Science | Prohibited | Yes | Not applicable | Prohibited | Jan 2023 |
| JAMA | Prohibited | Yes, tool + extent | Not applicable | Not explicitly addressed | 2023 |
| NEJM | Prohibited | Yes | Not applicable | Not explicitly addressed | 2023 |
| Cell Press | Prohibited | Yes, dedicated section | Not applicable | Not explicitly addressed | 2023 |
| NIH (grants) | Not applicable | Guidance only, not a submission requirement | Applications substantially developed by AI are not treated as the applicant’s original ideas and will not be considered; six applications per PI per calendar year (NOT-OD-25-132) | Prohibited (NOT-OD-23-149) | Jun 2023 / Jul 2025 |
| NSF (grants) | Not applicable | Encouraged, not required | Not stated | Prohibited: reviewers may not upload proposal content to non-approved tools | Dec 2023 |
7.6 Research Integrity Risks
Two integrity risks in AI-assisted research deserve explicit attention: citation hallucination and data fabrication.
Citation hallucination — the generation of plausible-sounding but non-existent references — is the most widely documented AI failure mode in research contexts. In the systematic review evaluation described earlier, hallucinated citations ranged from 28.6% of GPT-4’s output to 91.4% of Bard’s (Chelli et al. 2024). A study that asked ChatGPT to produce short literature reviews on 42 topics found that 55% of GPT-3.5’s citations and 18% of GPT-4’s referred to papers that do not exist, and among the citations that did correspond to real papers, 43% and 24% respectively contained substantive errors (Walters and Wilder 2023). The pattern is recognizable: a hallucinated citation typically has a plausible author list, a plausible journal, a plausible year, and a DOI that either does not exist or resolves to a different paper. The same fluency defeats human screening elsewhere in the manuscript. Blinded reviewers failed to flag roughly a third of AI-generated medical abstracts as machine-written, and those abstracts passed plagiarism detectors with near-perfect originality scores (Gao et al. 2023).
An investigator who does not verify AI-generated citations before submission will eventually submit a manuscript with fabricated references. Whether that becomes a misconduct finding turns on a standard worth stating precisely, because institutions write policy against it and quote it back at each other. Federal regulation defines research misconduct as fabrication, falsification, or plagiarism, and expressly excludes honest error and differences of opinion. A finding additionally requires a significant departure from accepted practice, conduct committed intentionally, knowingly, or recklessly, and proof by a preponderance of the evidence (U.S. Department of Health and Human Services 2024). An honest citation error is therefore not misconduct, and an institutional policy claiming otherwise will not survive contact with a research integrity officer. Submitting citations the investigator never checked, in a field that now knows exactly what these tools do, is a strong candidate for recklessness, which sits inside the standard. That is the narrower claim and the more defensible one.
The institutional response is a disclosure and verification norm: any AI tool used for literature-related tasks must have its output verified before it enters a manuscript or grant application. This is not a burden unique to AI — investigators are expected to have read the papers they cite regardless of how they found them. AI makes it easier to accumulate unchecked citations, so explicit verification practice is necessary.
Data fabrication through AI is a more serious concern that has attracted less attention in the policy literature. An investigator who asks an AI model to “fill in” missing data points, “smooth” irregularities in a dataset, or “generate” plausible results for an underpowered study is committing data fabrication in exactly the same way as manual fabrication. The fact that the fabrication was AI-assisted does not reduce the culpability. Research integrity training programs should explicitly address AI-assisted fabrication as a recognized misconduct category, not only the traditional forms.
7.7 The Research Trainee Gap
Everything in this chapter so far addresses the independent investigator. Yet much of the daily work of an academic medical center’s research enterprise is done by people still learning how to do it. A first-year PhD student in a translational lab, or an MD-PhD student writing a first-author manuscript, is already using AI for literature synthesis, for analysis code, and for drafting, often under a supervisor with less hands-on AI experience than the student. In most institutions in 2026, nobody has told either of them what the standard is. The education resources in this book address health professions learners; the workforce chapter addresses clinical and administrative staff. Research trainees sit in the seam between them, and the seam is where the next generation’s research habits are being set.
The gap is documented, not hypothetical. The AAMC’s AI competency work, which the Training & Workforce Development chapter discusses as the emerging reference point for academic medicine, is scoped to undergraduate, graduate, and continuing medical education (Association of American Medical Colleges 2025). That work was still a draft in community-feedback phase when this chapter was written, with a final report expected in 2026, so the claim here is about the draft and should be checked against the final version. “Graduate” in that scope means graduate medical education, which is to say residency. PhD students, postdoctoral fellows, and biomedical research training programs appear nowhere in it. The pattern is not American: the United Kingdom’s Vitae framework, the reference researcher-development standard in the Anglophone world, was refreshed in 2025, and neither the domains nor the descriptors of the resulting RDF 2025 name AI at all (Vitae 2025). A 2026 policy report by the quantitative methodologist Michael Zyphur identified that omission (Zyphur 2026); the published framework confirms it. The frameworks that define what a trained researcher should know looked directly at the question and did not see it.
Zyphur’s report, published through Instats, a commercial research-training provider that he directs, is the most complete argument to date that universities have mis-filed the AI question, and its central observation survives its commercial context. Institutions that have addressed AI in research training at all have mostly done so through plagiarism policy, and that is a category error. Plagiarism rules exist to police the copying of other people’s words; AI-generated text resembles that offense only at the surface. Nobody treats SPSS output as plagiarism when it appears in a results section, and nobody calls a machine-transcribed interview fraudulent so long as the researcher verified it against the audio. The questions that matter about a research tool are whether its use was ethical, valid, reproducible, and transparent. The reframing carries a concrete institutional consequence, and it is not the one that first suggests itself. Plagiarism is not the student-conduct alternative to research misconduct; it is one of the three federal categories of it, alongside fabrication and falsification, so a trainee who plagiarizes in federally funded work already belongs to the research integrity officer. The real jurisdictional gap sits elsewhere. AI use that is undisclosed, unverified, or methodologically invalid, but that involves no fabrication, falsification, or plagiarism, falls outside the federal definition altogether. It is not federally cognizable misconduct. An institution that wants its research integrity officer to hold that conduct has to confer the jurisdiction by local policy, because nothing confers it automatically, and the office will otherwise decline a case it has no authority over. Writing that authority down is a specific thing an AMC can do this year, and most have not done it.
The report also documents how unevenly the research system has responded. Within roughly ten weeks of ChatGPT’s release the major publishers had converged on one rule, that AI cannot be an author because it cannot bear responsibility; Table 7.1 records the biomedical version of that convergence and the disclosure requirements that followed it. Funders moved later and less uniformly, from the DFG’s statement in September 2023 to Australian funder positions issued in April 2026. University research-training policy, in the report’s international survey of thirty-eight universities, mostly stopped at research integrity without reaching AI literacy or valid research practice. That last count deserves the weight of an attributed observation rather than an established fact, since the survey’s coding rules are the author’s own and its underlying evidence files have not been published for independent review.
The publisher datapoint does useful work on its own, though less than it first appears. Sector-scale convergence was achieved in a quarter, which relocates the institutional excuse rather than removing it: these questions are not too new to answer. They are harder than the publishers’ case makes them look. The publishers had ICMJE authorship criteria and COPE machinery already built, and needed only to point them at a new object. Designing a research-training competency framework means deciding what a trainee must be able to do, at what stage, assessed by whom, under supervisors who mostly cannot do it themselves. Speed on a well-posed question does not dissolve the difficulty of an ill-posed one. What it does remove is the claim that nobody knows where to start.
The disclosure numbers suggest the answers are needed. Across 25,114 empirical research submissions to 49 BMJ Group journals over seven months of 2024, a mandatory structured disclosure field recorded AI use in 5.7 percent of them, against survey estimates that between 28 and 76 percent of researchers use AI in their work (AlFayyad et al. 2026). That gap is not by itself proof of concealment. A new field has adoption lag, and disclosure did climb from 4.5 to 7.3 percent across the period. Researchers also disagree about what counts as disclosable use, and improving the quality of one’s own writing, which 87 percent of the disclosing authors reported doing, is exactly the case where a working scientist might reasonably think no declaration is owed. That last possibility is the one an institution should find most actionable, because it is the one a competency framework fixes.
What should trainees actually be taught? The competencies translate readily into biomedical practice. Verify every citation before it enters a manuscript, as discussed above. Report the model, version, and parameters used in an analysis the way a methods section reports a reagent lot or an instrument setting. Treat prompt selection as a researcher degree of freedom: the analytic-flexibility work that gave biomedicine the concept of p-hacking (Simmons et al. 2011) applies with full force to AI-assisted analysis, and so do its remedies, preregistering prompts alongside analysis plans and testing whether results survive prompt variation. Two further habits are worth teaching as practitioner heuristics rather than as findings, because the evaluation literature behind them is moving too fast to cite with confidence. Run critique of a draft through a different model family than the one that produced it, on the assumption that a model reviewing its own output will tend to endorse it. And treat a model’s agreement as a sycophancy signal rather than a validation signal, a habit that transfers directly to how a trainee should weigh any flattering feedback on their science.
Preregistering a prompt is harder than preregistering an analysis, and the chapter should not pretend otherwise. A statistical analysis plan is verifiable because the analysis is deterministic given the data. The same prompt against the same nominal model can return different output across calls, across provider-side updates that carry no version change, and across model versions that are simply retired and cannot be re-run at all. Until that settles, the honest institutional position is that prompt preregistration establishes what was intended rather than what is replayable, and that the gateway log described below is what makes the intention checkable.
None of this requires an institution to invent standards from scratch. A statement in PNAS from a group convened by the National Academy of Sciences together with the Annenberg Public Policy Center and the Annenberg Foundation Trust at Sunnylands, whose authors include Marcia McNutt, Eric Horvitz, and Barbara Grosz, sets out five principles for protecting scientific integrity in the age of generative AI: transparent disclosure and attribution, verification of AI-generated content and analyses, documentation of AI-generated data, attention to ethics and equity, and continuous monitoring, oversight, and public engagement (Blau et al. 2024). An AMC can adopt those five principles as the spine of its research-training expectations and let supervision norms, methods-section requirements, and integrity-office adjudication hang from them. Project 3 at the end of this chapter is that adoption written as a project with an owner and a deadline.
The infrastructure this chapter recommends does double duty here. The secure AI gateway of Project 2 logs model, version, prompt, and parameters by user and project. That log is the beginning of a reproducibility record rather than the whole of one, and the distinction is worth drawing, because naming what is missing turns a general appeal to logging into a list an institution can put in a contract. Logging the input preserves what was asked. Getting the same answer back additionally requires capturing the output alongside the request, retaining whatever retrieval context grounded it, pinning to a model version the provider still serves, and keeping all of it longer than a misconduct inquiry takes, which is measured in years against commercial model lifecycles measured in months. For the integrity use the log also has to be tamper-evident and held outside the researcher’s control, which is a design property worth specifying before procurement rather than discovering after an allegation. The section that follows argues for the same gateway from HIPAA obligations; the research-training case arrives at it from reproducibility, and two independent arguments reaching the same conclusion make a stronger procurement case than either alone. The binding constraint, though, is not infrastructure. The report’s argument, and it is the right one, is that a candidate cannot acquire competencies the supervisor does not model. The workforce chapter documents the faculty development gap in clinical AI education; this is the same gap seen from the research side, and an AMC that funds faculty development for clinical AI literacy but not for research supervision has solved half of the problem — the half that does not end up in the published record.
7.8 Human Subjects, Privacy, and Secure Infrastructure
Research involving human subjects data requires particular care in AI tool selection and use. The Common Rule governs research involving identifiable private information; HIPAA governs the use and disclosure of protected health information. Neither was written with AI in mind, and institutional interpretation of how they apply to AI tool use is still evolving.
The operative question for an investigator is: what data can be entered into which AI tool? A reasonable institutional framework distinguishes four scenarios.
First, publicly available data or fully de-identified data (under HIPAA Safe Harbor or Expert Determination) may generally be entered into enterprise AI tools — provided the institutional BAA covers the tool. Whether it can be entered into consumer AI tools (public API, no BAA) depends on whether the data remains genuinely non-identifiable at the population level. The re-identification literature is clear that combinations of demographic variables that appear non-identifying individually can uniquely identify individuals in sparse datasets. When in doubt, treat data from clinical sources as identifiable.
Second, limited dataset (HIPAA limited dataset, which retains dates and geographic data at the county level) requires a data use agreement and should be treated as potentially identifiable for AI tool purposes. Enterprise tools with BAAs are appropriate; public APIs are not.
Third, the full EHR — identifiable clinical data — requires a BAA with the AI tool provider and should not be entered into any tool that does not have that agreement in place. Research use of full EHR data through AI tools should go through the institution’s IRB and through the honest broker process if the investigator is not the treating provider.
Fourth, genomic and genetic data is subject to additional constraints — the Genetic Information Nondiscrimination Act (GINA), NIH data sharing policies for GWAS data, and the special re-identification risk of genomic data that is well documented in the literature.
The infrastructure conclusion from this framework is that the research enterprise needs institutional API access to a language model that operates within an enterprise tenant with BAA, not just general permission to use consumer tools. This is a capital and contract decision that should go through the AI Steering Committee, not an investigator-by-investigator determination.
A distinct data governance issue that has emerged from NIH’s 2023 Data Management and Sharing Policy is how model weights and training datasets from AI research projects are treated as “scientific data” subject to sharing requirements. In March 2025 NIH answered part of that question for genomics. Notice NOT-OD-25-081 treats generative AI models trained on NIH controlled-access genomic data, model parameters included, as Data Derivatives under the Genomic Data Sharing Policy and the Data Use Certification. Such a model may not be shared outside the Approved User group, must not be retained after the project closes, and controlled-access data may not be submitted to public generative AI tools at all (National Institutes of Health 2025a). An investigator who trains a model on dbGaP data and then posts the weights has made a data-sharing disclosure, not a software release. The practical implications — what must be shared, what may be shared, and what cannot be shared under DMS obligations — are addressed in Chapter 17.
Figure 7.1 illustrates the full research lifecycle with AI tool availability and constraints annotated at each stage.
7.9 Where to Start: Three Starter Projects
7.9.1 Project 1: Institutional Literature Review Toolkit
What it is. Negotiate an institutional subscription to Elicit or a comparable semantic search tool and configure it with access for all investigators. Create a one-page usage guide that specifies: which tasks the tool is appropriate for (scoping reviews, rapid evidence maps, initial literature exploration), which tasks require supplementation with structured database search (systematic reviews intended for peer review, Cochrane-style meta-analyses), and how to document AI tool use in methods sections per target journal requirements.
What you need to start. A research computing or library liaison who can manage the institutional subscription, budget for an annual license, and someone in the research integrity or library team who can write the one-page guide. The guide should take one working day to draft and one meeting with a few senior investigators to validate.
How you will know it works here. Before the first renewal, run a one-page acceptance test: take three completed systematic reviews from your own investigators, in the fields your institution actually publishes in, and measure what fraction of each final reference list the tool surfaces from the original research question. A psychiatry corpus and a structural biology corpus are not the same retrieval problem, and vendor recall figures are averaged over neither. The book argues elsewhere that benchmark performance does not transfer to local deployment; that principle does not stop applying because the tool is a literature search rather than a clinical model.
Build or buy? Buy — specifically, subscribe to an existing tool. Do not build a literature search AI from scratch. The commercial tools have invested years in recall optimization over biomedical corpora; a locally built vector search over PubMed will not match them and will require ongoing infrastructure investment to maintain.
What done looks like. Within 90 days: institutional account active, usage guide published on the research computing or library website, at least 20 investigators have used the tool, and a short usage report has been delivered to the research dean showing which departments are using it and for what. The usage data is itself valuable: it shows where investigators are investing in literature work and where training is needed.
7.9.2 Project 2: Secure AI Gateway for Research Computing
What it is. Deploy an institutional API gateway that routes requests to a language model (GPT-4, Claude, or Gemini) through an enterprise-tenanted endpoint with a BAA, logs all usage by user and project, and enforces data classification rules — specifically, rejecting or flagging requests that contain patterns matching PHI or genomic identifiers. Researchers submit queries through a simple web interface or through an API key issued against their institutional credentials.
Why this matters for research specifically. The alternative — investigators using personal OpenAI or Anthropic accounts for research tasks — creates two problems. It routes research data through personal accounts with no institutional oversight or logging. It makes it impossible to answer the question “what AI tools did you use on this project?” when a journal or funder asks, because there is no institutional record.
What you need to start. The core technical components are an Azure OpenAI or AWS Bedrock deployment (both offer BAAs; setup takes one to two weeks for a system administrator familiar with the platform), a simple API gateway (LiteLLM proxy is an open-source option that adds routing, logging, and cost tracking with minimal configuration), and a credential issuance process tied to the institution’s identity provider. The entire system can be built by one competent platform engineer in two to four weeks.
Why investigators will actually use it. A funded PI can charge API access to a grant without asking anyone, so the personal account is not a policy violation to be closed but a legitimate procurement the investigator already controls. A gateway that trails the frontier by a model generation, adds latency, or filters prompts opaquely will lose that comparison on the merits and should. The design goal is therefore to make the institutional path the easier one: no procurement paperwork, no personal card, higher rate limits than an individual can buy, current models, and the logging as a benefit the investigator wants rather than a tax they pay. If the gateway cannot win on convenience, adoption numbers will tell you so, and mandates will not fix it.
Build or buy? Build the gateway, buy the model. The model capability comes from the commercial provider; the gateway is lightweight infrastructure that the institution controls. This is one of the cases where a modest engineering investment produces a durable capability: once the gateway is running, adding new model providers, new data classification rules, or new usage analytics is incremental work.
What done looks like. Investigators can make API requests to the institutional gateway from their preferred coding environment (R, Python, Jupyter) or from a simple web interface. Usage is logged by project and researcher. A data classification filter is running on prompt content. A quarterly usage report goes to the research dean and the AISC. When a journal asks “did you use AI on this project?”, the investigator can answer based on their logged requests rather than on memory.
7.9.3 Project 3: Research Trainee AI Standard and Supervisor Development
What it is. A short statement of what a research trainee at this institution is expected to do when using AI, a matching expectation for the people supervising them, and an amendment to the research misconduct policy that gives the research integrity officer jurisdiction over AI use falling outside fabrication, falsification, and plagiarism. The statement takes the five PNAS principles as its spine (Blau et al. 2024) and translates each into something a second-year PhD student can actually check on a Friday afternoon: verify every citation against the source before it enters a manuscript, report model and version and parameters in the methods section, preregister prompts alongside the analysis plan, run critique through a different model family than the one that drafted the text, and treat a model’s agreement as a sycophancy signal.
Why this matters for research specifically. The other two projects buy tools. This one sets the habits that determine whether the tools produce a defensible record, and it is the only one that addresses the population actually doing most of the bench and analytic work. It is also the institution’s answer to the enforcement problem in the NIH originality rule: a trainee accused of submitting AI-developed text is in a far better position if the institution can show what standard they were taught and what their gateway log contains.
What you need to start. The research integrity officer, the graduate school or postdoctoral affairs office, and two or three program directors from PhD and MD-PhD programs. The policy amendment needs whoever owns the research misconduct policy, which is usually the same office that houses the RIO. Budget a week of one person’s time for the drafting and a full semester for the amendment, because policy amendments move at the speed of the committee that owns them rather than the speed of the drafting.
Build or buy? Build, though not from a blank page. The five principles are published, the competencies are set out earlier in this chapter, and peer institutions are drafting the same document right now and will generally share a draft if asked. What cannot be bought is the jurisdiction amendment and the supervisor component, both of which depend on how this particular institution is governed and who reports to whom.
What done looks like. Within one academic year: the trainee standard is published and referenced in the onboarding of every PhD and postdoctoral appointment; graduate students and postdoctoral fellows hold credentials on the Project 2 gateway in their own names rather than borrowing a supervisor’s; the research misconduct policy names AI misuse outside the federal categories and says who adjudicates it; every research training program has run at least one supervisor session on the same material, funded on the same budget line as clinical AI faculty development; and the research integrity officer can say how many AI-related allegations arrived that year and how they were resolved. The last of those is what tells you whether any of the rest of it is real.