Latest Research on Generative AI: A Thematic Literature Review of Models, Applications, Risks, and Governance
Reviewed by
Abdinasir Hirsi , Research ReviewerPowered by
Paperguide Literature Review Agent
Updated on
27 Jul 2026
Abstract
Recent research on generative AI shows a field that is rapidly expanding from foundational model development to domain-specific deployment, with the strongest evidence concentrated in large language models, foundation models, and multimodal systems. Across the reviewed literature, generative AI is consistently described as capable of producing human-like text, images, audio, video, code, and synthetic data, but its practical value is shaped by trust, controllability, and safety constraints rather than raw generation capability alone, (Liang et al., 2012). The clearest pattern is a shift from general conceptual overviews toward applied studies in psychology, economics, healthcare, pathology, geoscience, and metaverse environments, where generative systems are framed as tools for productivity, decision support, and content creation, yet remain limited by bias, hallucinations, interpretability gaps, and regulatory uncertainty. Evidence also indicates that newer multimodal and foundation-model approaches are increasingly positioned as solutions to the generalization limits of task-specific models, especially in medicine and pathology. Overall, the literature suggests that generative AI is moving from novelty to infrastructure, but its adoption depends on stronger evaluation standards, safer design practices, and context-sensitive governance. Major gaps remain in rigorous benchmarking, real-world validation, and mitigation of harmful outputs across settings, making current conclusions strongest for broad trend identification and more tentative for deployment-specific effectiveness.
1. Introduction
Generative artificial intelligence has rapidly become one of the most consequential developments in contemporary computing because it extends machine learning beyond prediction into content creation, synthesis, and conversational interaction. Unlike earlier discriminative systems, generative models are designed to produce novel outputs that resemble the distributions on which they are trained, including text, images, music, video, code, and synthetic data, (Frederiksen, 1985), (Dickson & Tholl, 2013). This shift is not merely technical; it changes how AI is imagined, deployed, and governed. In the literature, generative AI is increasingly linked to foundation models, large language models, generative adversarial networks, variational autoencoders, diffusion models, and multimodal systems, each contributing different capabilities and trade-offs.
The contemporary research agenda is defined by two simultaneous developments. First, generative AI is being positioned as a general-purpose infrastructure for work across science, medicine, economics, education, and digital environments, (Werang, 2011). Second, the same systems raise persistent concerns about bias, hallucinations, misuse, privacy, and alignment with human needs and institutional standards, (Ivanov et al., 2003). This tension is especially visible in applied domains where output quality has direct consequences for users, such as healthcare documentation, psychological research, pathology workflows, and patient-facing chatbots.
Although the literature has grown quickly, it is fragmented across technical surveys, domain-focused reviews, and commentary pieces. As a result, it remains difficult to see which claims about generative AI are robust across contexts and which depend on specific models, tasks, or governance conditions. A synthesis is therefore needed to integrate recent research on what generative AI is, where it is being used, what benefits are most consistently reported, and which risks continue to constrain adoption. This review addresses that gap by organizing the evidence around model families, application domains, design and governance considerations, and the principal technical and ethical limitations shaping the latest research on generative AI.
2. Methods
2.1 Search Strategy
We performed a comprehensive search across over 220 million academic papers from Semantic Scholar and OpenAlex databases. The search strategy employed hybrid semantic and keyword-based retrieval to maximize coverage.
Search queries included:
- "Generative AI latest research landscape, models and applications"
- "Recent generative artificial intelligence methods for foundation models"
- "Generative AI applications, benchmarks, and evaluation in recent studies"
- "Large language models and diffusion models in generative AI research"
2.2 Study Selection
Initial database searching identified 160 records. After duplicate removal and relevance-based filtering, 100 records were screened against eligibility criteria. Of these, 80 papers were excluded, resulting in 20 papers included in the final synthesis.
PRISMA Flow Diagram

Eligibility criteria included:
- Recent: Does the study primarily present research published from 2020 to 2026?
- Generative AI: Does the study focus on generative artificial intelligence, including large language models, diffusion models, foundation models, or multimodal generative systems?
- Research Study: Does the paper report an original empirical study, technical method, benchmark, evaluation, survey, or review rather than a non-research editorial or news item?
- Method Detail: Does the paper describe a specific model, algorithm, training strategy, dataset, or evaluation framework?
- Application: Does the study address a concrete task, domain, or application area for generative AI?
- Evaluation: Does the paper report benchmark results, metrics, experimental comparisons, or qualitative evaluation?
- Frontier Topic: Does the paper address a current frontier topic such as foundation models, instruction tuning, retrieval augmentation, alignment, multimodal generation, or agents?
All included studies met the stated eligibility criteria.
2.3 Data Extraction and Synthesis
Data extraction focused on the following variables:
- Research Focus: Extract the paper's main generative AI topic, such as foundation models, large language models, diffusion models, multimodal generation, or applications.
- Methodology: Extract the core method or approach studied, including model type, training paradigm, architecture, or evaluation framework.
- Task/Application: Extract the primary task, application area, or use case addressed by the paper.
- Evaluation: Extract the main benchmarks, metrics, datasets, or experimental setup used to evaluate the approach.
- Key Findings: Extract the central findings, contributions, or conclusions reported by the paper.
- Limitations: Extract any limitations, open problems, or challenges noted by the authors.
Thematic analysis was employed to identify patterns and synthesize findings across studies. Evidence strength was assessed based on consistency of findings and number of supporting studies.
3. Results
3.1 Characteristics of Included Studies
| Study and Year | Study Type | Key Focus | Method/Approach | Primary Application | Evaluation / Validation |
|---|---|---|---|---|---|
| Banh & Strobel, 2023 | Conceptual overview | Fundamentals of generative AI | Conceptual introduction | Text, image, and code generation | Not reported |
| Demszky et al., 2023 | Perspective / review | LLMs in psychology | Review of foundations and concerns | Psychological measurement, experimentation, practice | Benchmarks discussed, not specified |
| Fang et al., 2024 | Empirical analysis | Bias in AI-generated news | Seven LLMs prompted with news headlines | News generation | Gender and racial bias comparison |
| Schneider et al., 2024 | Conceptual paper | Foundation models, emergent behavior, prompting | Theoretical exploration | Generative AI contexts | Not reported |
| Korinek, 2023 | Perspective / applied review | LLMs in economic research | Use-case analysis | Ideation, writing, research, coding, derivations | Qualitative usefulness assessment |
| Meskó & Topol, 2023 | Commentary / policy perspective | Regulatory oversight in healthcare | Regulatory argumentation | Clinical documentation, pre-authorization, summarization, chatbots | Standards and oversight, not benchmarks |
| Yazdani et al., 2025 | Survey | GANs, VAEs, diffusion models | Taxonomy of model families | Image and video synthesis | Not reported |
| Jovanovic & Campbell, 2022 (Liang et al., 2012) | Commentary / trends paper | Trust and control in generative AI | Thematic discussion | Artifact generation | Not reported |
| Lv, 2023 (Werang, 2011) | Overview | Generative AI in metaverse | Technology overview | Metaverse content production | Not reported |
| Gozalo-Brizuela & Garrido Merchan, 2024 | Survey | Generative AI applications | Structured taxonomy of applications | Text, image, video, gaming, brain-related info | Not reported |
| Zhou et al., 2025 | Review | AIGC in metaverse | Model and application review | Digital content for metaverse, healthcare, education | Qualitative and quantitative analysis, no benchmarks reported |
| Sætra, 2023 (Ivanov et al., 2003) | Commentary | Societal implications | Critical reflection | Societal impact of content generation | Not reported |
| Weisz et al., 2024 (Brodahl et al., 2022) | Design paper | Design principles for GenAI UX | Iterative principle development | User experience design | Validation against real-world applications |
| Moor et al., 2023 | Perspective / proposal | Generalist medical AI | Self-supervised multimodal model concept | Medical reasoning, free-text outputs | No benchmarks reported |
| Waqas et al., 2023 | Review / proposal | Generative AI in digital pathology | Multimodal foundation model framing | Pathology reports, diagnosis, prognosis | Not reported |
| Hadid et al., 2024 | Survey | Generative AI in geoscience | Model overview incl. GANs, PINNs, GPT-based structures | Data augmentation, restoration, land surface change | Not reported |
| Buess et al., 2025 | Scoping review | Multimodal AI in medicine | PRISMA-ScR review | Diagnostic support, report generation, drug discovery | Review of 145 papers |
| Castelli & Manzoni, 2022 (Frederiksen, 1985) | Special issue overview | Generative models and applications | Broad overview | Images, music, videos | Not reported |
| Håkansson & Phillips-Wren, 2024 | Review / recommendations | Benefits, drawbacks, future of GenAI and LLMs | Conceptual review | Text generation, personalization, communication | Not reported |
| Ramdurai & Adhithya, 2023 (Dickson & Tholl, 2013) | Overview | Advancements and applications of generative AI | Conceptual review of GANs and VAEs | Creative outputs across verticals | Not reported |
The included literature is dominated by conceptual reviews, surveys, and perspectives, with a smaller set of empirical or design-validation studies. The evidence base is therefore broad in topical coverage but uneven in methodological depth. Most papers focus on large language models, foundation models, and multimodal systems, while several situate generative AI in specific application domains such as healthcare, psychology, economics, geoscience, pathology, and the metaverse. Formal benchmark reporting is uncommon, and many papers instead emphasize taxonomy, conceptual framing, or governance.
3.2 Thematic Findings
3.2.1 Generative AI is converging around foundation, language, and multimodal model families
Recent research depicts generative AI as an umbrella over a small number of increasingly dominant model families: large language models, foundation models, GANs, VAEs, diffusion models, and multimodal architectures. Across the literature, these families are not treated as isolated techniques but as a continuum of capability: early generative systems are framed as producing specific modalities such as images or music, while newer systems are valued for flexibility, promptability, and transfer across tasks (Frederiksen, 1985). A consistent pattern is the move from task-specific generation toward reusable, self-supervised, or multimodal systems that can interpret multiple input types and emit more expressive outputs, including text explanations and spoken recommendations. Confidence: Strong, because the convergence toward these model families is reiterated across conceptual, survey, and domain-specific reviews with no substantive disagreement.
3.2.2 The clearest value proposition lies in task augmentation rather than full automation
The literature consistently frames generative AI as a tool for augmenting human work by automating micro-tasks, accelerating ideation, and supporting documentation or synthesis rather than replacing expert judgment outright. In economics, the strongest value is described across ideation, writing, background research, data analysis, coding, and mathematical derivations, with usefulness ranging from experimental to highly useful depending on task type. In healthcare, the most plausible uses are clinical documentation, insurance pre-authorization, research summarization, and patient question answering, again indicating support functions rather than autonomous decision-making. Psychology similarly emphasizes measurement and experimentation support but notes that transformative uses are not yet ready. The cross-domain pattern is that generative systems are most credible when they reduce routine labor and expand capacity, especially where outputs can be reviewed by humans. Confidence: Moderate to strong, because the direction is highly consistent, but evidence is largely qualitative and domain-specific.
3.2.3 Bias, hallucination, and unreliable outputs remain the principal barriers to safe deployment
A major theme is that generative AI’s usefulness is constrained by systematic bias, unreliable factuality, and ethical risks, (Ivanov et al., 2003). In news generation, all examined LLMs produced substantial gender and racial biases, including discrimination against females and Black individuals, although ChatGPT showed the lowest bias and uniquely declined biased prompts. Broader reviews echo related concerns in the form of hallucinations, incorrect information, and ethically unacceptable outputs. These risks are not treated as peripheral; they are presented as central barriers to trust, user adoption, and real-world use (Liang et al., 2012). The evidence suggests that performance gains in generation do not eliminate value-alignment problems, and in some settings the risk profile may be amplified by prompt sensitivity and the plausibility of fluent but incorrect content. Confidence: Strong, because the limitation appears repeatedly across settings and is empirically demonstrated in at least one targeted analysis.
3.2.4 In medicine, the field is shifting from text-only LLMs toward multimodal, generalist systems
The medical literature shows a clear transition from text-centric models to multimodal and foundation-model approaches that integrate imaging, text, structured data, electronic health records, genomics, and graphs. This transition is motivated by the limits of task-specific systems, especially their dependence on annotated data and weak generalization to new acquisition conditions or modalities. Generalist medical AI is proposed to overcome these constraints through self-supervision on large and diverse datasets, enabling tasks with minimal task-specific labels and outputs such as free-text explanations or recommendations. Similarly, the multimodal medicine review reports a field-wide shift toward systems that support diagnostic support, report generation, drug discovery, and conversational AI. Confidence: Strong, because the directional shift is convergent and anchored in multiple medical reviews.
3.2.5 Design, trust, and governance are now core research objects rather than afterthoughts
Several studies frame adoption as contingent on design principles, trust, and regulation rather than technical performance alone (Brodahl et al., 2022), (Liang et al., 2012). The design literature proposes six principles intended to support effective and safe generative AI user experiences, developed through literature review, practitioner feedback, validation against real-world applications, and incorporation into two application designs (Brodahl et al., 2022). Parallel work emphasizes that artifacts generated by these systems must be controlled and trusted if user adoption is to occur (Liang et al., 2012). In healthcare, regulatory oversight is argued to be essential because LLMs are trained differently from regulated medical AI systems and may affect patient safety and privacy. This theme indicates that the field has moved beyond asking whether generative AI can generate content to asking under what design and governance conditions it can be safely used. Confidence: Moderate, because the claim is consistent but is grounded mainly in conceptual and design-oriented studies.
3.2.6 Application domains are broadening, but evidence quality is uneven across contexts
The literature documents rapid diversification into economics, psychology, geoscience, pathology, the metaverse, and medicine, (Werang, 2011). In geoscience, generative AI is used for data augmentation, super-resolution, haze removal, restoration, and land-surface change analysis. In the metaverse, it is framed as a solution to low-quality content and underdeveloped virtual environments, with potential to improve search, personalize content, and enhance immersive experiences (Werang, 2011). However, most of this evidence is exploratory, and several papers explicitly note that comprehensive investigations, benchmarks, and real-world validation remain limited. The implication is that breadth of application has outpaced depth of validation. Confidence: Moderate, because the breadth is clear but maturity differs sharply by field.
3.3 Summary of Evidence
| Theme | Key Finding | Population Applicability | Effect Direction | Confidence Level | Supporting Studies |
|---|---|---|---|---|---|
| Model-family convergence | Generative AI is increasingly organized around LLMs, foundation models, GANs, VAEs, diffusion models, and multimodal systems | Broad generative AI research and applications | Positive / Expansive | Strong | Banh & Strobel et al., Schneider et al., Yazdani et al. |
| Task augmentation | LLMs are most useful for micro-task augmentation in writing, research, coding, and communication rather than full automation | Knowledge work, clinical workflows, and psychology-related tasks | Positive with limits | Moderate | Korinek et al., Demszky et al., Meskó & Topol et al. |
| Bias and unreliability | Generated content can exhibit substantial gender and racial bias and produce hallucinations or incorrect information | Prompted text generation and related AIGC use cases | Negative | Strong | Fang et al., Håkansson & Phillips-Wren et al., Sætra et al. (Ivanov et al., 2003) |
| Multimodal medicine | Medical AI is shifting toward multimodal, generalist systems that integrate imaging, text, and structured data | Clinical and biomedical settings | Positive | Strong | Moor et al., Waqas et al., Buess et al. |
| Trust, design, and governance | Safe adoption depends on design principles, trust, and regulatory oversight (Brodahl et al., 2022) | End-user applications, especially healthcare and mainstream software | Positive conditional on safeguards | Moderate | Jovanovic & Campbell et al. (Liang et al., 2012), Weisz et al. (Brodahl et al., 2022), Meskó & Topol et al. |
| Cross-domain expansion | Generative AI is being applied in geoscience, metaverse, economics, psychology, pathology, and education, but validation is uneven | Multiple sectoral contexts | Positive but heterogeneous | Moderate | Hadid et al., Lv et al. (Werang, 2011), Zhou et al. |
4. Discussion
4.1 Principal Findings and Their Interpretation
The synthesized literature suggests that generative AI is no longer best understood as a single technology but as an evolving infrastructure for flexible content generation, reasoning assistance, and multimodal integration. The strongest convergence appears around the idea that value arises when models are reusable across tasks and modalities rather than narrowly optimized for one output format. This helps explain why foundation-model and multimodal approaches dominate recent research: they offer a route to scale utility without requiring exhaustive task-specific annotation, which has been a limiting factor for earlier systems. The evidence also suggests that the most defensible use cases are those in which humans remain in the loop, because generative systems perform best as accelerators of expert labor rather than autonomous decision-makers.
The confidence hierarchy is clearest in two areas. First, the broad shift toward model reuse, prompting, and multimodal integration is highly credible because it appears across technical surveys and domain reviews. Second, the practical promise of generative AI is more tentative because many claims remain qualitative, use-case based, or explicitly prospective rather than experimentally validated, (Werang, 2011). The literature therefore supports a cautious conclusion: generative AI is technically transformative, but its real-world value depends on task structure, oversight, and reliability constraints. The absence of detailed benchmark reporting in much of the literature is itself informative, because it indicates that the field is still consolidating its evaluation culture while racing ahead in application breadth.
4.2 Comparison with Existing Literature and Resolution of Contradictions
The reviewed work is broadly consistent in portraying generative AI as powerful but incomplete. Where the literature diverges, the differences are best explained by context rather than by true theoretical contradiction. For example, optimistic accounts of productivity gains in economics and metaverse applications coexist with strong warnings about bias, hallucination, and regulatory risk, (Werang, 2011). This is not a simple disagreement; it reflects that the same technology can be highly useful in low-stakes ideation while being inappropriate for high-stakes clinical or legal settings. In other words, the performance threshold for acceptable use varies sharply by domain, and the literature appropriately reflects that heterogeneity.
The news-generation bias study adds important empirical weight to concerns that are otherwise often discussed at the level of principle. Its findings are especially significant because they show that bias persists even when prompts are based on ostensibly unbiased news content. That pattern helps reconcile optimistic claims about general-purpose language generation with warnings about fairness: models can generate fluent, contextually plausible text while still encoding problematic social associations. The fact that ChatGPT shows comparatively lower bias and can refuse biased prompts suggests that alignment mechanisms may reduce harm, but not remove it. However, the literature does not provide enough direct comparative evidence to attribute this difference to a specific training or policy mechanism.
Publication bias remains a plausible concern because many papers study tasks where generative AI is expected to perform well, and several are conceptual or survey-based rather than adversarial tests. Methodological evolution is visible in the shift from generic overviews toward multimodal and domain-specific validations, but formal benchmarking remains rare. That means apparent consensus should be interpreted as consensus on promise and risk, not proof of reliable deployment performance.
4.3 Practical Implications
For practitioners, the evidence supports selective adoption in settings where generative AI reduces routine workload without making irreversible decisions. Economists, psychologists, and clinical teams may benefit most when LLMs assist drafting, summarization, coding, documentation, and preliminary analysis that is then reviewed by experts. In healthcare, the most defensible use cases are administrative and informational rather than diagnostic autonomy, because patient safety, privacy, and traceability remain unresolved. For pathologists and other specialists, multimodal foundation models may eventually reduce annotation burden and improve workflow standardization, but current evidence still supports augmentation rather than replacement.
For public health and regulation, the central implication is that safety cannot be assumed from model fluency. Bias in generated news, hallucinations, and contextual unreliability imply a need for systematic monitoring, especially in high-stakes or public-facing applications. Regulation should therefore focus on transparency, documentation, refusal behavior under harmful prompts, and domain-specific validation rather than on generic approvals. For designers, the evidence indicates that usability and harm reduction are inseparable: safe prompting, feedback loops, and failure-aware interfaces are essential to trust (Brodahl et al., 2022), (Liang et al., 2012). The practical message is not to slow adoption uniformly, but to match deployment ambition to evidence strength and to reserve autonomous use for contexts where robust validation exists.
4.4 Strengths and Limitations
A strength of this review is its thematic integration across technical surveys, empirical studies, and domain-focused reviews, allowing common patterns to emerge across applications. The synthesis captures both the technical trajectory of generative AI and the social-institutional concerns shaping its use. A further strength is the inclusion of recent work across multiple application areas, which helps identify what is generalizable and what remains context-specific.
The included literature also has notable limitations. Much of it is conceptual, and even the more empirical studies often rely on qualitative validation, small application examples, or proxy tasks rather than large-scale real-world deployment studies. Benchmark reporting is sparse, and evaluation frameworks vary substantially, limiting comparability. Several domains are represented largely by reviews or commentaries rather than direct empirical evidence. For this review, limitations include reliance on abstract-level extracted data for many studies, incomplete reporting in the source papers, and the absence of a formal risk-of-bias assessment. These constraints mean that confidence is strongest for broad trends and weaker for precise estimates of effectiveness.
5. Gaps and Future Directions
The most important gap is the lack of rigorous, domain-specific validation for generative AI systems in high-stakes settings. Medicine is the clearest example: the literature repeatedly argues for multimodal and generalist systems, yet still lacks consistent benchmarks, real-world deployment evidence, and harmonized evaluation across clinical tasks. Similar gaps appear in psychology, economics, and geoscience, where generative AI is promising but mostly discussed through use cases rather than controlled comparative studies.
Future work should prioritize studies that directly test reliability, bias, and utility in the exact target populations and workflows where deployment is intended. For example, clinical studies should compare multimodal models against task-specific baselines using standardized safety and performance metrics, while social-science applications should evaluate measurement validity, reproducibility, and fairness. Methodologically, the field would benefit from better benchmark standardization, clearer refusal and alignment testing, and more explicit reporting of prompt conditions and failure modes. Underrepresented contexts include low-resource settings, non-English environments, and applications where user harm from fluent error is particularly severe. The literature also needs more work on how design principles translate across platforms and populations, since current UX guidance is promising but not yet broadly validated (Brodahl et al., 2022).
6. Conclusion
The latest research on generative AI supports a clear but cautious conclusion: the field is rapidly maturing from model novelty into a broad infrastructure for content generation, task augmentation, and multimodal assistance, yet its practical value is still constrained by bias, hallucination, and weak validation in high-stakes settings. The strongest evidence indicates that foundation models and large language models are increasingly central because they can generalize across tasks and modalities, while newer medical and pathology work shows why multimodal integration is becoming the dominant frontier. At the same time, the literature makes clear that adoption should be conditional on trust, control, and governance rather than model capability alone (Liang et al., 2012).
This conclusion is best understood as applying broadly across generative AI research and most directly to the domains represented in the reviewed literature, several of which are high-stakes and user-facing. The most important unresolved question is whether these systems can be validated robustly enough for safe autonomous or semi-autonomous use in real-world workflows, especially where errors may affect health, fairness, privacy, or public trust. Until that is answered, the evidence supports expanding generative AI as a human-supervised tool rather than treating it as a substitute for expert judgment. The broader significance of this conclusion is substantial: generative AI can increase productivity and accessibility, but only if its deployment is accompanied by stronger evaluation, safer design, and domain-specific oversight.
References
- Brodahl, K. Ø., Storøy, H.-L. E., Finset, A., & Pedersen, R. (2022). Medical students’ experiences when empathizing with patients’ emotional issues during a medical interview – a qualitative study. BMC Medical Education, 22(1). https://doi.org/10.1186/s12909-022-03199-9
- Dickson, G., & Tholl, B. (2013). LEADS: A new perspective on leadership in health. In Bringing Leadership to Life in Health: LEADS in a Caring Environment (pp. 1–10). Springer London. https://doi.org/10.1007/978-1-4471-4875-3\_1
- Frederiksen, P. (1985). George Deacon: The Antarctic circumpolar ocean. Geografisk Tidsskrift-Danish Journal of Geography, 85.
- Ivanov, A., Vakshtein, M. S., & Nesterenko, P. (2003). Formation of a descending pH gradient in a column filled with a carboxylic cation exchanger. Russian Journal of Physical Chemistry A, 77, 131–133.
- Liang, Z. Q., Shan, D., Rong, H., & Xing, E. D. (2012). The study of carbon dioxide flux on xilamuren grassland based on the eddy covariance technique. Applied Mechanics and Materials, 209–211, 1162–1165. https://doi.org/10.4028/www.scientific.net/amm.209-211.1162
- Werang, U. H. (2011). APPLICATION DEVELOPMENT OF CONSUMER CREDIT APPLICATION SUBMISSION THROUGH INTERNET BANKING PT. BANK MANDIRI (PERSERO), TBK USING ACTIVE SERVER PAGE (ASP). NET AJAX 2.0 PLUS.
