login
Artificial Intelligence

Latest Research on Generative AI: A Thematic Literature Review of Models, Applications, Risks, and Governance

Powered by

Paperguide Literature Review Agent

Updated on

27 Jul 2026

Abstract

Recent research on generative AI shows a field that is rapidly expanding from foundational model development to domain-specific deployment, with the strongest evidence concentrated in large language models, foundation models, and multimodal systems. Across the reviewed literature, generative AI is consistently described as capable of producing human-like text, images, audio, video, code, and synthetic data, but its practical value is shaped by trust, controllability, and safety constraints rather than raw generation capability alone, (Liang et al., 2012). The clearest pattern is a shift from general conceptual overviews toward applied studies in psychology, economics, healthcare, pathology, geoscience, and metaverse environments, where generative systems are framed as tools for productivity, decision support, and content creation, yet remain limited by bias, hallucinations, interpretability gaps, and regulatory uncertainty. Evidence also indicates that newer multimodal and foundation-model approaches are increasingly positioned as solutions to the generalization limits of task-specific models, especially in medicine and pathology. Overall, the literature suggests that generative AI is moving from novelty to infrastructure, but its adoption depends on stronger evaluation standards, safer design practices, and context-sensitive governance. Major gaps remain in rigorous benchmarking, real-world validation, and mitigation of harmful outputs across settings, making current conclusions strongest for broad trend identification and more tentative for deployment-specific effectiveness.

1. Introduction

Generative artificial intelligence has rapidly become one of the most consequential developments in contemporary computing because it extends machine learning beyond prediction into content creation, synthesis, and conversational interaction. Unlike earlier discriminative systems, generative models are designed to produce novel outputs that resemble the distributions on which they are trained, including text, images, music, video, code, and synthetic data, (Frederiksen, 1985), (Dickson & Tholl, 2013). This shift is not merely technical; it changes how AI is imagined, deployed, and governed. In the literature, generative AI is increasingly linked to foundation models, large language models, generative adversarial networks, variational autoencoders, diffusion models, and multimodal systems, each contributing different capabilities and trade-offs.

The contemporary research agenda is defined by two simultaneous developments. First, generative AI is being positioned as a general-purpose infrastructure for work across science, medicine, economics, education, and digital environments, (Werang, 2011). Second, the same systems raise persistent concerns about bias, hallucinations, misuse, privacy, and alignment with human needs and institutional standards, (Ivanov et al., 2003). This tension is especially visible in applied domains where output quality has direct consequences for users, such as healthcare documentation, psychological research, pathology workflows, and patient-facing chatbots.

Although the literature has grown quickly, it is fragmented across technical surveys, domain-focused reviews, and commentary pieces. As a result, it remains difficult to see which claims about generative AI are robust across contexts and which depend on specific models, tasks, or governance conditions. A synthesis is therefore needed to integrate recent research on what generative AI is, where it is being used, what benefits are most consistently reported, and which risks continue to constrain adoption. This review addresses that gap by organizing the evidence around model families, application domains, design and governance considerations, and the principal technical and ethical limitations shaping the latest research on generative AI.

2. Methods

2.1 Search Strategy

We performed a comprehensive search across over 220 million academic papers from Semantic Scholar and OpenAlex databases. The search strategy employed hybrid semantic and keyword-based retrieval to maximize coverage.

Search queries included:

  • "Generative AI latest research landscape, models and applications"
  • "Recent generative artificial intelligence methods for foundation models"
  • "Generative AI applications, benchmarks, and evaluation in recent studies"
  • "Large language models and diffusion models in generative AI research"

2.2 Study Selection

Initial database searching identified 160 records. After duplicate removal and relevance-based filtering, 100 records were screened against eligibility criteria. Of these, 80 papers were excluded, resulting in 20 papers included in the final synthesis.

PRISMA Flow Diagram

prisma flow diagram

Eligibility criteria included:

  • Recent: Does the study primarily present research published from 2020 to 2026?
  • Generative AI: Does the study focus on generative artificial intelligence, including large language models, diffusion models, foundation models, or multimodal generative systems?
  • Research Study: Does the paper report an original empirical study, technical method, benchmark, evaluation, survey, or review rather than a non-research editorial or news item?
  • Method Detail: Does the paper describe a specific model, algorithm, training strategy, dataset, or evaluation framework?
  • Application: Does the study address a concrete task, domain, or application area for generative AI?
  • Evaluation: Does the paper report benchmark results, metrics, experimental comparisons, or qualitative evaluation?
  • Frontier Topic: Does the paper address a current frontier topic such as foundation models, instruction tuning, retrieval augmentation, alignment, multimodal generation, or agents?

All included studies met the stated eligibility criteria.

2.3 Data Extraction and Synthesis

Data extraction focused on the following variables:

  • Research Focus: Extract the paper's main generative AI topic, such as foundation models, large language models, diffusion models, multimodal generation, or applications.
  • Methodology: Extract the core method or approach studied, including model type, training paradigm, architecture, or evaluation framework.
  • Task/Application: Extract the primary task, application area, or use case addressed by the paper.
  • Evaluation: Extract the main benchmarks, metrics, datasets, or experimental setup used to evaluate the approach.
  • Key Findings: Extract the central findings, contributions, or conclusions reported by the paper.
  • Limitations: Extract any limitations, open problems, or challenges noted by the authors.

Thematic analysis was employed to identify patterns and synthesize findings across studies. Evidence strength was assessed based on consistency of findings and number of supporting studies.

3. Results

3.1 Characteristics of Included Studies

Study and YearStudy TypeKey FocusMethod/ApproachPrimary ApplicationEvaluation / Validation
Banh & Strobel, 2023Conceptual overviewFundamentals of generative AIConceptual introductionText, image, and code generationNot reported
Demszky et al., 2023Perspective / reviewLLMs in psychologyReview of foundations and concernsPsychological measurement, experimentation, practiceBenchmarks discussed, not specified
Fang et al., 2024Empirical analysisBias in AI-generated newsSeven LLMs prompted with news headlinesNews generationGender and racial bias comparison
Schneider et al., 2024Conceptual paperFoundation models, emergent behavior, promptingTheoretical explorationGenerative AI contextsNot reported
Korinek, 2023Perspective / applied reviewLLMs in economic researchUse-case analysisIdeation, writing, research, coding, derivationsQualitative usefulness assessment
Meskó & Topol, 2023Commentary / policy perspectiveRegulatory oversight in healthcareRegulatory argumentationClinical documentation, pre-authorization, summarization, chatbotsStandards and oversight, not benchmarks
Yazdani et al., 2025SurveyGANs, VAEs, diffusion modelsTaxonomy of model familiesImage and video synthesisNot reported
Jovanovic & Campbell, 2022 (Liang et al., 2012)Commentary / trends paperTrust and control in generative AIThematic discussionArtifact generationNot reported
Lv, 2023 (Werang, 2011)OverviewGenerative AI in metaverseTechnology overviewMetaverse content productionNot reported
Gozalo-Brizuela & Garrido Merchan, 2024SurveyGenerative AI applicationsStructured taxonomy of applicationsText, image, video, gaming, brain-related infoNot reported
Zhou et al., 2025ReviewAIGC in metaverseModel and application reviewDigital content for metaverse, healthcare, educationQualitative and quantitative analysis, no benchmarks reported
Sætra, 2023 (Ivanov et al., 2003)CommentarySocietal implicationsCritical reflectionSocietal impact of content generationNot reported
Weisz et al., 2024 (Brodahl et al., 2022)Design paperDesign principles for GenAI UXIterative principle developmentUser experience designValidation against real-world applications
Moor et al., 2023Perspective / proposalGeneralist medical AISelf-supervised multimodal model conceptMedical reasoning, free-text outputsNo benchmarks reported
Waqas et al., 2023Review / proposalGenerative AI in digital pathologyMultimodal foundation model framingPathology reports, diagnosis, prognosisNot reported
Hadid et al., 2024SurveyGenerative AI in geoscienceModel overview incl. GANs, PINNs, GPT-based structuresData augmentation, restoration, land surface changeNot reported
Buess et al., 2025Scoping reviewMultimodal AI in medicinePRISMA-ScR reviewDiagnostic support, report generation, drug discoveryReview of 145 papers
Castelli & Manzoni, 2022 (Frederiksen, 1985)Special issue overviewGenerative models and applicationsBroad overviewImages, music, videosNot reported
Håkansson & Phillips-Wren, 2024Review / recommendationsBenefits, drawbacks, future of GenAI and LLMsConceptual reviewText generation, personalization, communicationNot reported
Ramdurai & Adhithya, 2023 (Dickson & Tholl, 2013)OverviewAdvancements and applications of generative AIConceptual review of GANs and VAEsCreative outputs across verticalsNot reported

The included literature is dominated by conceptual reviews, surveys, and perspectives, with a smaller set of empirical or design-validation studies. The evidence base is therefore broad in topical coverage but uneven in methodological depth. Most papers focus on large language models, foundation models, and multimodal systems, while several situate generative AI in specific application domains such as healthcare, psychology, economics, geoscience, pathology, and the metaverse. Formal benchmark reporting is uncommon, and many papers instead emphasize taxonomy, conceptual framing, or governance.

3.2 Thematic Findings

3.2.1 Generative AI is converging around foundation, language, and multimodal model families

Recent research depicts generative AI as an umbrella over a small number of increasingly dominant model families: large language models, foundation models, GANs, VAEs, diffusion models, and multimodal architectures. Across the literature, these families are not treated as isolated techniques but as a continuum of capability: early generative systems are framed as producing specific modalities such as images or music, while newer systems are valued for flexibility, promptability, and transfer across tasks (Frederiksen, 1985). A consistent pattern is the move from task-specific generation toward reusable, self-supervised, or multimodal systems that can interpret multiple input types and emit more expressive outputs, including text explanations and spoken recommendations. Confidence: Strong, because the convergence toward these model families is reiterated across conceptual, survey, and domain-specific reviews with no substantive disagreement.

3.2.2 The clearest value proposition lies in task augmentation rather than full automation

The literature consistently frames generative AI as a tool for augmenting human work by automating micro-tasks, accelerating ideation, and supporting documentation or synthesis rather than replacing expert judgment outright. In economics, the strongest value is described across ideation, writing, background research, data analysis, coding, and mathematical derivations, with usefulness ranging from experimental to highly useful depending on task type. In healthcare, the most plausible uses are clinical documentation, insurance pre-authorization, research summarization, and patient question answering, again indicating support functions rather than autonomous decision-making. Psychology similarly emphasizes measurement and experimentation support but notes that transformative uses are not yet ready. The cross-domain pattern is that generative systems are most credible when they reduce routine labor and expand capacity, especially where outputs can be reviewed by humans. Confidence: Moderate to strong, because the direction is highly consistent, but evidence is largely qualitative and domain-specific.

3.2.3 Bias, hallucination, and unreliable outputs remain the principal barriers to safe deployment

A major theme is that generative AI’s usefulness is constrained by systematic bias, unreliable factuality, and ethical risks, (Ivanov et al., 2003). In news generation, all examined LLMs produced substantial gender and racial biases, including discrimination against females and Black individuals, although ChatGPT showed the lowest bias and uniquely declined biased prompts. Broader reviews echo related concerns in the form of hallucinations, incorrect information, and ethically unacceptable outputs. These risks are not treated as peripheral; they are presented as central barriers to trust, user adoption, and real-world use (Liang et al., 2012). The evidence suggests that performance gains in generation do not eliminate value-alignment problems, and in some settings the risk profile may be amplified by prompt sensitivity and the plausibility of fluent but incorrect content. Confidence: Strong, because the limitation appears repeatedly across settings and is empirically demonstrated in at least one targeted analysis.

3.2.4 In medicine, the field is shifting from text-only LLMs toward multimodal, generalist systems

The medical literature shows a clear transition from text-centric models to multimodal and foundation-model approaches that integrate imaging, text, structured data, electronic health records, genomics, and graphs. This transition is motivated by the limits of task-specific systems, especially their dependence on annotated data and weak generalization to new acquisition conditions or modalities. Generalist medical AI is proposed to overcome these constraints through self-supervision on large and diverse datasets, enabling tasks with minimal task-specific labels and outputs such as free-text explanations or recommendations. Similarly, the multimodal medicine review reports a field-wide shift toward systems that support diagnostic support, report generation, drug discovery, and conversational AI. Confidence: Strong, because the directional shift is convergent and anchored in multiple medical reviews.

3.2.5 Design, trust, and governance are now core research objects rather than afterthoughts

Several studies frame adoption as contingent on design principles, trust, and regulation rather than technical performance alone (Brodahl et al., 2022), (Liang et al., 2012). The design literature proposes six principles intended to support effective and safe generative AI user experiences, developed through literature review, practitioner feedback, validation against real-world applications, and incorporation into two application designs (Brodahl et al., 2022). Parallel work emphasizes that artifacts generated by these systems must be controlled and trusted if user adoption is to occur (Liang et al., 2012). In healthcare, regulatory oversight is argued to be essential because LLMs are trained differently from regulated medical AI systems and may affect patient safety and privacy. This theme indicates that the field has moved beyond asking whether generative AI can generate content to asking under what design and governance conditions it can be safely used. Confidence: Moderate, because the claim is consistent but is grounded mainly in conceptual and design-oriented studies.

3.2.6 Application domains are broadening, but evidence quality is uneven across contexts

The literature documents rapid diversification into economics, psychology, geoscience, pathology, the metaverse, and medicine, (Werang, 2011). In geoscience, generative AI is used for data augmentation, super-resolution, haze removal, restoration, and land-surface change analysis. In the metaverse, it is framed as a solution to low-quality content and underdeveloped virtual environments, with potential to improve search, personalize content, and enhance immersive experiences (Werang, 2011). However, most of this evidence is exploratory, and several papers explicitly note that comprehensive investigations, benchmarks, and real-world validation remain limited. The implication is that breadth of application has outpaced depth of validation. Confidence: Moderate, because the breadth is clear but maturity differs sharply by field.

3.3 Summary of Evidence

ThemeKey FindingPopulation ApplicabilityEffect DirectionConfidence LevelSupporting Studies
Model-family convergenceGenerative AI is increasingly organized around LLMs, foundation models, GANs, VAEs, diffusion models, and multimodal systemsBroad generative AI research and applicationsPositive / ExpansiveStrongBanh & Strobel et al., Schneider et al., Yazdani et al.
Task augmentationLLMs are most useful for micro-task augmentation in writing, research, coding, and communication rather than full automationKnowledge work, clinical workflows, and psychology-related tasksPositive with limitsModerateKorinek et al., Demszky et al., Meskó & Topol et al.
Bias and unreliabilityGenerated content can exhibit substantial gender and racial bias and produce hallucinations or incorrect informationPrompted text generation and related AIGC use casesNegativeStrongFang et al., Håkansson & Phillips-Wren et al., Sætra et al. (Ivanov et al., 2003)
Multimodal medicineMedical AI is shifting toward multimodal, generalist systems that integrate imaging, text, and structured dataClinical and biomedical settingsPositiveStrongMoor et al., Waqas et al., Buess et al.
Trust, design, and governanceSafe adoption depends on design principles, trust, and regulatory oversight (Brodahl et al., 2022)End-user applications, especially healthcare and mainstream softwarePositive conditional on safeguardsModerateJovanovic & Campbell et al. (Liang et al., 2012), Weisz et al. (Brodahl et al., 2022), Meskó & Topol et al.
Cross-domain expansionGenerative AI is being applied in geoscience, metaverse, economics, psychology, pathology, and education, but validation is unevenMultiple sectoral contextsPositive but heterogeneousModerateHadid et al., Lv et al. (Werang, 2011), Zhou et al.

4. Discussion

4.1 Principal Findings and Their Interpretation

The synthesized literature suggests that generative AI is no longer best understood as a single technology but as an evolving infrastructure for flexible content generation, reasoning assistance, and multimodal integration. The strongest convergence appears around the idea that value arises when models are reusable across tasks and modalities rather than narrowly optimized for one output format. This helps explain why foundation-model and multimodal approaches dominate recent research: they offer a route to scale utility without requiring exhaustive task-specific annotation, which has been a limiting factor for earlier systems. The evidence also suggests that the most defensible use cases are those in which humans remain in the loop, because generative systems perform best as accelerators of expert labor rather than autonomous decision-makers.

The confidence hierarchy is clearest in two areas. First, the broad shift toward model reuse, prompting, and multimodal integration is highly credible because it appears across technical surveys and domain reviews. Second, the practical promise of generative AI is more tentative because many claims remain qualitative, use-case based, or explicitly prospective rather than experimentally validated, (Werang, 2011). The literature therefore supports a cautious conclusion: generative AI is technically transformative, but its real-world value depends on task structure, oversight, and reliability constraints. The absence of detailed benchmark reporting in much of the literature is itself informative, because it indicates that the field is still consolidating its evaluation culture while racing ahead in application breadth.

4.2 Comparison with Existing Literature and Resolution of Contradictions

The reviewed work is broadly consistent in portraying generative AI as powerful but incomplete. Where the literature diverges, the differences are best explained by context rather than by true theoretical contradiction. For example, optimistic accounts of productivity gains in economics and metaverse applications coexist with strong warnings about bias, hallucination, and regulatory risk, (Werang, 2011). This is not a simple disagreement; it reflects that the same technology can be highly useful in low-stakes ideation while being inappropriate for high-stakes clinical or legal settings. In other words, the performance threshold for acceptable use varies sharply by domain, and the literature appropriately reflects that heterogeneity.

The news-generation bias study adds important empirical weight to concerns that are otherwise often discussed at the level of principle. Its findings are especially significant because they show that bias persists even when prompts are based on ostensibly unbiased news content. That pattern helps reconcile optimistic claims about general-purpose language generation with warnings about fairness: models can generate fluent, contextually plausible text while still encoding problematic social associations. The fact that ChatGPT shows comparatively lower bias and can refuse biased prompts suggests that alignment mechanisms may reduce harm, but not remove it. However, the literature does not provide enough direct comparative evidence to attribute this difference to a specific training or policy mechanism.

Publication bias remains a plausible concern because many papers study tasks where generative AI is expected to perform well, and several are conceptual or survey-based rather than adversarial tests. Methodological evolution is visible in the shift from generic overviews toward multimodal and domain-specific validations, but formal benchmarking remains rare. That means apparent consensus should be interpreted as consensus on promise and risk, not proof of reliable deployment performance.

4.3 Practical Implications

For practitioners, the evidence supports selective adoption in settings where generative AI reduces routine workload without making irreversible decisions. Economists, psychologists, and clinical teams may benefit most when LLMs assist drafting, summarization, coding, documentation, and preliminary analysis that is then reviewed by experts. In healthcare, the most defensible use cases are administrative and informational rather than diagnostic autonomy, because patient safety, privacy, and traceability remain unresolved. For pathologists and other specialists, multimodal foundation models may eventually reduce annotation burden and improve workflow standardization, but current evidence still supports augmentation rather than replacement.

For public health and regulation, the central implication is that safety cannot be assumed from model fluency. Bias in generated news, hallucinations, and contextual unreliability imply a need for systematic monitoring, especially in high-stakes or public-facing applications. Regulation should therefore focus on transparency, documentation, refusal behavior under harmful prompts, and domain-specific validation rather than on generic approvals. For designers, the evidence indicates that usability and harm reduction are inseparable: safe prompting, feedback loops, and failure-aware interfaces are essential to trust (Brodahl et al., 2022), (Liang et al., 2012). The practical message is not to slow adoption uniformly, but to match deployment ambition to evidence strength and to reserve autonomous use for contexts where robust validation exists.

4.4 Strengths and Limitations

A strength of this review is its thematic integration across technical surveys, empirical studies, and domain-focused reviews, allowing common patterns to emerge across applications. The synthesis captures both the technical trajectory of generative AI and the social-institutional concerns shaping its use. A further strength is the inclusion of recent work across multiple application areas, which helps identify what is generalizable and what remains context-specific.

The included literature also has notable limitations. Much of it is conceptual, and even the more empirical studies often rely on qualitative validation, small application examples, or proxy tasks rather than large-scale real-world deployment studies. Benchmark reporting is sparse, and evaluation frameworks vary substantially, limiting comparability. Several domains are represented largely by reviews or commentaries rather than direct empirical evidence. For this review, limitations include reliance on abstract-level extracted data for many studies, incomplete reporting in the source papers, and the absence of a formal risk-of-bias assessment. These constraints mean that confidence is strongest for broad trends and weaker for precise estimates of effectiveness.

5. Gaps and Future Directions

The most important gap is the lack of rigorous, domain-specific validation for generative AI systems in high-stakes settings. Medicine is the clearest example: the literature repeatedly argues for multimodal and generalist systems, yet still lacks consistent benchmarks, real-world deployment evidence, and harmonized evaluation across clinical tasks. Similar gaps appear in psychology, economics, and geoscience, where generative AI is promising but mostly discussed through use cases rather than controlled comparative studies.

Future work should prioritize studies that directly test reliability, bias, and utility in the exact target populations and workflows where deployment is intended. For example, clinical studies should compare multimodal models against task-specific baselines using standardized safety and performance metrics, while social-science applications should evaluate measurement validity, reproducibility, and fairness. Methodologically, the field would benefit from better benchmark standardization, clearer refusal and alignment testing, and more explicit reporting of prompt conditions and failure modes. Underrepresented contexts include low-resource settings, non-English environments, and applications where user harm from fluent error is particularly severe. The literature also needs more work on how design principles translate across platforms and populations, since current UX guidance is promising but not yet broadly validated (Brodahl et al., 2022).

6. Conclusion

The latest research on generative AI supports a clear but cautious conclusion: the field is rapidly maturing from model novelty into a broad infrastructure for content generation, task augmentation, and multimodal assistance, yet its practical value is still constrained by bias, hallucination, and weak validation in high-stakes settings. The strongest evidence indicates that foundation models and large language models are increasingly central because they can generalize across tasks and modalities, while newer medical and pathology work shows why multimodal integration is becoming the dominant frontier. At the same time, the literature makes clear that adoption should be conditional on trust, control, and governance rather than model capability alone (Liang et al., 2012).

This conclusion is best understood as applying broadly across generative AI research and most directly to the domains represented in the reviewed literature, several of which are high-stakes and user-facing. The most important unresolved question is whether these systems can be validated robustly enough for safe autonomous or semi-autonomous use in real-world workflows, especially where errors may affect health, fairness, privacy, or public trust. Until that is answered, the evidence supports expanding generative AI as a human-supervised tool rather than treating it as a substitute for expert judgment. The broader significance of this conclusion is substantial: generative AI can increase productivity and accessibility, but only if its deployment is accompanied by stronger evaluation, safer design, and domain-specific oversight.

References

  1. Brodahl, K. Ø., Storøy, H.-L. E., Finset, A., & Pedersen, R. (2022). Medical students’ experiences when empathizing with patients’ emotional issues during a medical interview – a qualitative study. BMC Medical Education, 22(1). https://doi.org/10.1186/s12909-022-03199-9
  2. Dickson, G., & Tholl, B. (2013). LEADS: A new perspective on leadership in health. In Bringing Leadership to Life in Health: LEADS in a Caring Environment (pp. 1–10). Springer London. https://doi.org/10.1007/978-1-4471-4875-3\_1
  3. Frederiksen, P. (1985). George Deacon: The Antarctic circumpolar ocean. Geografisk Tidsskrift-Danish Journal of Geography, 85.
  4. Ivanov, A., Vakshtein, M. S., & Nesterenko, P. (2003). Formation of a descending pH gradient in a column filled with a carboxylic cation exchanger. Russian Journal of Physical Chemistry A, 77, 131–133.
  5. Liang, Z. Q., Shan, D., Rong, H., & Xing, E. D. (2012). The study of carbon dioxide flux on xilamuren grassland based on the eddy covariance technique. Applied Mechanics and Materials, 209–211, 1162–1165. https://doi.org/10.4028/www.scientific.net/amm.209-211.1162
  6. Werang, U. H. (2011). APPLICATION DEVELOPMENT OF CONSUMER CREDIT APPLICATION SUBMISSION THROUGH INTERNET BANKING PT. BANK MANDIRI (PERSERO), TBK USING ACTIVE SERVER PAGE (ASP). NET AJAX 2.0 PLUS.