


Review
paper
The future of pharmaceuticals: Arti
fi
cial intelligence in drug discovery
and
development
Chen
Fu
a
,
b
,
Qiuchen
Chen
a
,
c
,
*
a
Department
of
Pharmacology,
School
of
Pharmacy,
China
Medical
University,
Shenyang, 110122,
China
b
Pharmaceutical
Sciences
Laboratory
Center,
School
of
Pharmacy,
China
Medical
University,
Shenyang, 110122,
China
c
Liaoning
Key
Laboratory
of
Molecular
Targeted
Anti-tumor
Drug
Development
and
Evaluation,
China
Medical
University,
Shenyang, 110122,
China
a
r
t
i
c
l
e
i
n
f
o
Article
history:
Received
30
July
2024
Received
in
revised
form
13
February
2025
Accepted
21
February
2025
Available
online
26
February
2025
Keywords:
AI
Drugs
Research
and
development
Machine
learning
a
b
s
t
r
a
c
t
Arti
fi
cial
Intelligence
(AI)
is
revolutionizing
traditional
drug
discovery
and
development
models
by
seamlessly integrating data, computational power, and algorithms. This synergy enhances the ef
fi
ciency,
accuracy, and success rates of drug research, shortens development timelines, and reduces costs. Coupled
with
machine
learning
(ML)
and
deep
learning
(DL),
AI
has
demonstrated
signi
fi
cant
advancements
across various
domains, including
drug characterization,
target discovery and validation,
small molecule
drug
design,
and
the
acceleration
of
clinical
trials.
Through
molecular
generation
techniques,
AI
facili-
tates
the
creation
of
novel
drug
molecules,
predicting
their
properties
and
activities,
while
virtual
screening (VS) optimizes drug candidates. Additionally, AI enhances clinical trial ef
fi
ciency by predicting
outcomes,
designing
trials,
and
enabling
drug
repositioning.
However,
AI's
application
in
drug
devel-
opment faces challenges, including the need for robust data-sharing mechanisms and the establishment
of
more
comprehensive
intellectual
property
protections
for
algorithms.
AI-driven
pharmaceutical
companies
must
also
integrate
biological
sciences
and
algorithms
effectively,
ensuring
the
successful
fusion
of
wet
and
dry
laboratory
experiments.
Despite
these
challenges,
the
potential
of
AI
in
drug
development
remains
undeniable.
As
AI
technology
evolves
and
these
barriers
are
addressed,
AI-driven
therapeutics
are
poised
for
a
broader
and
more
impactful
future
in
the
pharmaceutical
industry.
©
2025
The
Author(s).
Published
by
Elsevier
B.V.
on
behalf
of
Xi
’
an
Jiaotong
University.
This
is
an
open
access
article
under
the
CC
BY-NC-ND
license
(
http://creativecommons.org/licenses/by-nc-nd/4.0/
).
1.
Introduction
AI,
de
fi
ned
as
the
intelligence
demonstrated
by
human-made
machines,
emerged
as
a
novel
fi
eld
of
study
dedicated
to
devel-
oping
theories,
methods,
technologies,
and
applications
aimed
at
simulating, extending, and enhancing human intelligence [
1
]. Over
the past six decades, AI has evolved from a theoretical concept into
a
powerful
industrial
tool,
revolutionizing
industries
such
as
manufacturing,
agriculture,
healthcare,
and
fi
nance
[
2
e
5
].
AI
technologies
have
been
successfully
applied
in
areas
including
autonomous
driving,
voice
recognition,
web
search,
and
medical
diagnosis
[
6
,
7
].
Its
capabilities
in
specialized
tasks
like
language
translation
and
facial
recognition
now
rival
or
surpass
human
performance, leading to the observation that
“
no
fi
eld is immune to
the
charms
and
sweep
of
AI".
A watershed
moment
in
AI's
history
occurred in March 2016 when AlphaGo, an AI program, triumphed
over
renowned
South
Korean
Go
player
Lee
Sedol,
sparking
wide-
spread societal debate [
8
]. The 2024 Nobel Prize in Physics went to
two
scientists,
John
J.
Hop
fi
eld
and
Geoffrey
E.
Hinton,
for
foun-
dational
discoveries
and
inventions
that
enable
machine
learning
with
arti
fi
cial
neural
networks.
Furthermore,
the
Nobel
Prize
in
Chemistry
recognized
using
AI
to
design proteins.
Proteins are
the
workhorse
molecules
of
life,
with
millions
existing
in
nature,
but
novel
ones
could
transform
medicine
and
technology.
The
new
tools
have
already enabled
researchers
to
churn
out
designer
pro-
teins
for
vaccines
and
cancer
treatment,
arti
fi
cial
pollution-eating
enzymes, and molecular assemblies capable of seeding the growth
of
minerals
[
9
,
10
].
2.
The
history
of
the
development
of
AI
drugs
AI has been applied in the pharmaceutical
fi
eld for nearly three
decades,
with
signi
fi
cant
advancements
since
the
late
1990s
in
its
underlying
algorithmic
frameworks.
Over
this
period,
AI
has
un-
dergone periods of
rapid
progress and
setbacks
[
11
], driven by the
evolution
from
neural
networks
to
deep
neural
networks
(DNNs)
Peer
review
under
responsibility
of
Xi'an
Jiaotong
University.
*
Corresponding author. Department of Pharmacology, School of Pharmacy, China
Medical
University,
Shenyang, 110122,
China.
E-mail
address:
qcchen@cmu.edu.cn
(Q.
Chen).
Contents
lists
available
at
ScienceDirect
Journal
of
Pharmaceutical
Analysis
journal
homepage:
www.elsevier.com/locate/jpa
https://doi.org/10.1016/j.jpha.2025.101248
2095-1779/
©
2025
The
Author(s).
Published
by
Elsevier
B.V.
on
behalf
of
Xi
’
an
Jiaotong
University.
This
is
an
open
access
article
under
the
CC
BY-NC-ND
license
(
http://
creativecommons.org/licenses/by-nc-nd/4.0/
).
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248

and from ML to DL. Continuous optimization of algorithms, coupled
with
growing
data
accumulation
and
computational
power,
has
been instrumental in advancing the AI
fi
eld [
12
,
13
]. To evaluate the
absorption,
distribution,
metabolism,
excretion,
and
toxicity
(ADMET)
properties
of
new
molecular
entities
(NMEs)
as
early
as
possible,
various
in
vitro
and
in
vivo
methods
including
medium-
and
high-throughput
screening
have
been
developed,
which
also
facilitate the rapid accumulation of experimental data. However, as
the
number
of
NMEs
continues
to
increase,
these
experimental
approaches
have
shown
several
inherent
shortcomings:
time-
consuming, costly, and animal welfare issues involved, which have
greatly
limited
their
application
and
spurred
the
emergence
of
in
silico
methods for predicting ADMET properties. In recent decades,
with
the
rapid
development
of
computer
science
and
the
accu-
mulation
of
ADMET experimental
data,
in
silico
predictive
models
and derived web tools aimed at facilitating the ef
fi
cient evaluation
of
ADMET
properties
have
been
greatly
developed
[
14
].
Since
2018,
AI
in
pharmaceuticals
has
advanced
from
a
con-
ceptual
phase
(
“
0
”
)
to
practical
application
(
“
1
”
)
(
Fig. 1
).
Although
no AI-enabled drugs have been approved by the U.S. Food and Drug
Administration
(U.S.
FDA)
for
marketing
yet,
several
AI-driven
pharmaceutical
companies
have
successfully
accelerated
phases
I
and II clinical candidates [
15
]. In 2024, recent research highlighted
a
breakthrough
in
drug
design
using
DL
to
reverse-engineer
syn-
thetic
routes,
an
achievement
compared
to
AlphaGo's
impact
on
chemistry
[
16
,
17
].
This
marked
the
beginning
of
signi
fi
cant
break-
throughs in AI-driven pharmaceuticals. For instance, in 2021, Healx
utilized
AI
to
identify
new
uses
for
the
drug
HLX-0201
in
treating
fragile X syndrome, advancing the
project to phase II clinical
trials
within
18
months
[
18
].
In
2019,
Deep
Genomics
applied
its
AI
platform
to
identify
novel
targets
and
screen
oligonucleotide
can-
didates
for
Wilson's
disease,
completing
the
process
in
just
18
months
[
19
].
Insilico
Intelligence
(Insilico
Medicine)
utilized
GENTRL,
a
generative
adversarial
network
(GAN)-based
approach,
to
complete
an
AI
drug
discovery
challenge
within
21
days.
This
process
involved
data
collection,
model
development,
and
the
design
of
novel
molecules,
ultimately
generating
a
highly
active
discoidin domain receptor 1 (DDR1) kinase inhibitor [
20
]. Although
the
identi
fi
ed
compounds
demonstrated
satisfactory
microsomal
stability
and
pharmacokinetic
properties,
further
optimization
is
required
to
improve
selectivity,
speci
fi
city,
and
other
critical
medicinal
chemistry
parameters.
Additionally,
DeepMind's
Alpha-
Fold 3 achieved a signi
fi
cant breakthrough in addressing a 50-year-
old
biological
challenge
by
accurately
predicting
the
three-
dimensional
(3D)
structure
of
proteins
[
21
].
In
March
2024,
InSys
Intelligence's
fully
AI-generated
drug
for
idiopathic
pulmonary
fi
brosis
(IPF)
entered
phase
IIa
trials.
This
drug,
with
a
novel
backbone compound developed by Chemistry42 using AI software
Pandaomics,
showcases
AI's
potential
in
innovative
drug
develop-
ment [
22
]. However, it is crucial to recognize the limitations of AI in
drug discovery. Analyses derived from multiple AI methods may be
misleading due to issues like overlap between testing and training
datasets,
biases
in
the
data,
or
a
lack
of
chemical
insight
into
the
results. These biases can produce high apparent accuracy, but often
with
poor
generalizability
and
limited
applicability
in
prospective
research.
Nonetheless,
AI's
role
in
predicting
and
screening
new
thera-
peutic
targets
and
drugs
remains
a
highly
promising
area
of
research [
23
]. Over the years, AI has been consistently heralded as a
transformative
tool
in
accelerating
drug
discovery,
development,
and
testing,
signi
fi
cantly
reducing
research
timelines.
Presently,
numerous
AI-enabled
drug
development
pipelines
are
entering
clinical
phases
globally (
Table
1
)
[
24
e
64
].
3.
AI
pharmaceutical
elements
3.1.
The
elements
of
AI
The three core components of AI, i.e., data, computation, and al-
gorithms,
serve
as
the
foundation
of
AI-driven
pharmaceutical
research.
Data
sources
in
the
pharmaceutical
sector
include
public
and
commercial
datasets,
research
and
development
(R
&
D)
data
obtained through collaborations with pharmaceutical companies, in-
house
research
datasets,
and
those
generated
through
data
mining
and
manual
cleaning
and
validation.
Advancements
in
computa-
tional
power,
particularly
through
graphics
processing
unit
(GPU)
cloud
computing
resources,
have
provided
critical
support
for
AI
pharmaceutical
companies
[
65
].
Different
algorithmic
models
are
tailored
to
speci
fi
c
application
scenarios,
and
when
combined
with
unique
data
sources,
they
give
rise
to
the
distinctive
pro
fi
les
of
AI
companies [
66
].
Fig.
1.
Brief
overview
of
arti
fi
cial
intelligence
(AI)
pharmaceutical
development.
DL:
deep
learning;
CPU:
central
processing
unit;
GPU:
graphics
processing
unit;
NN:
neural
network.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
2
3.2.
AI,
ML,
and
DL
In
the
2020
CASP14
competition,
DeepMind's
AlphaFold
2
ach-
ieved
a
groundbreaking
advance
in
protein
structure
prediction,
outpacing
the
second-place
competitor
by
a
substantial
margin
[
67
]. Media outlets widely referred to this achievement using terms
such
as
AI,
ML,
and
DL
[
68
].
AI
encompasses
machine/computer
vision and NLP agents capable of perceiving their environment and
reacting to it to achieve speci
fi
c objectives (
Fig. 2
). The fundamental
approach involves
“
training
”
machines with algorithms and data to
enable
them
to
perform
tasks
and
make
predictions
or
inferences
about future outcomes. The industry often classi
fi
es ML algorithms
in
two
ways:
based
on
learning
scenarios
or
by
their
form
and
function
[
69
].
DL,
a
more
advanced
form
of
ML,
utilizes
combina-
torial non-linear models that automatically learn effective features
at
multiple
levels
from
high-dimensional,
complex
data
[
70
].
The
“
depth
”
in
DL
refers
to the
number of
layers
in
the
network;
more
layers indicate a deeper network. Each module within the network
transforms
its
input
into
higher-level,
more
abstract
representa-
tions
[
71
].
DL
models
are
often
considered
a
form
of
end-to-end
learning,
where
the
learning
process
is
integrated
without
sepa-
rating
it
into
discrete
modules
or
stages.
Instead,
“
input-output
”
pairs
are
provided
as
training
data,
which
the
system
optimizes
throughout
the
process
to
achieve
the
task
without
propagating
errors
common
in
traditional
ML
[
72
].
Currently,
DL
is
predomi-
nantly
based
on
neural
network
models,
which
feature
a
multi-
layered
structure
designed
to
mimic
the
human
brain
and
learn
from
large
data
volumes
[
73
].
These
systems
have
demonstrated
superior
accuracy
and
sophistication
compared
to
other
ML
methods,
proving
highly
successful
in
tasks
such
as
image
and
sound recognition [
74
]. DL has catalyzed AI's rapid growth over the
past
decade,
with
applications
already
deeply
embedded
in
com-
puter
vision
and
NLP.
Furthermore,
DL
is
emerging
as
a
powerful
tool
in
complex
fi
elds
like
high-energy
physics,
computational
chemistry,
computational
biology,
and
medical
diagnostics,
where
expert
knowledge
is
essential [
75
].
Previous
studies
have
demonstrated
that
DL
technology
offers
signi
fi
cant
advantages
in
optimizing
chemical
synthesis
routes,
predicting drug pharmacokinetics, identifying drug target sites, and
generating
new
molecular
structures
[
76
].
DL
models
learn
the
intrinsic
relationships
between
compounds
and
target
proteins
by
training
on
extensive
datasets
of
known
compound-target
protein
interactions.
This
training
enables
DL
models
to
automatically
extract relevant features from both compounds and target proteins,
as
well
as
discern
interaction
patterns.
For
example,
Guttman
and
Kerem [
77
] developed a cytochrome P-450 3A4 (CYP3A4) inhibitor
Table
1
Arti
fi
cial
Intelligence
(AI)-enabled
drugs
entering the
clinical
phase.
No.
Company
Pipeline
Indications
Clinical
phase
Clinical
trials
No.
Update
year
Refs.
1
Recursion
REC4881
Familial
adenomatous
polyposis
Phase
I
NCT05552755
2025
[
24
]
2
REC2282
Neuro
fi
bromatosis
type
2
Phase
II/III
NCT05130866
2024
[
25
]
3
REC994
Cerebral
cavernous
malformation
Phase
II
NCT05085561
2024
[
26
]
4
Lantern
LP100
mCRPC
Phase
II
NCT03643107
2025
[
27
]
5
LP300
Lung
adenocarcinoma
Phase
II
NCT05456256
2025
[
28
]
6
LP184
Solid
tumors
Phase
I
NCT05933265
2025
[
29
]
7
Relay
RLY1971
SHP2
Phase
I
NCT04252339
2023
[
30
]
8
RLY4008
FGFR2
Phase
I
NCT04526106
2025
[
31
]
9
RLY2608
PI3K
a
Phase
I
NCT05216432
2025
[
32
]
10
Accutar
Biotechnology
AC682
Breast
cancer
Terminated
NCT05489679
2024
[
33
]
11
AC176
mCRPC
Phase
I
NCT05241613
2025
[
34
]
12
Berg
Health
BPM31510
Glioblastoma
Phase
II
NCT04752813
2025
[
35
]
13
BPM31543
Alopecia
Phase
I
NCT01588522
2017
[
36
]
14
AI
Therapeutics
LAM-001
BOS
and
pulmonary
sarcoidosis
Phase
II
NCT05798923
2024
[
37
]
15
LAM-002A
Amyotrophic
lateral
sclerosis
Phase
II
NCT05163886
2024
[
38
]
16
Benevolent
AI
BEN-2293
Atopic
dermatitis
Phase
II
NCT04737304
2023
[
39
]
17
BioXcel
Therapeutics
BXCL501
Acute
agitation
Phase
II
NCT05276830
2023
[
40
]
18
BXCL701
Metastatic
castration-resistant
prostate
cancer
Phase
II
NCT03910660
2023
[
41
]
19
Exscientia
EXS21546
Oncology
Phase
I
NCT04727138
2022
[
42
]
20
Evaxion
Biotech
EXV-01
Metastatic
melanoma
Phase
II
NCT05309421
2023
[
43
]
21
EXV-02
Adjuvant
melanoma
Phase
I/II
NCT04455503
2024
[
44
]
22
Pharos
Ibio
PHI-101-001
Acute
myelogenous
leukemia
Phase
I
NCT04842370
2021
[
45
]
23
PHI-101-002
Platinum-resistant
ovarian
cancer
Phase
I
NCT04678102
2023
[
46
]
24
SOM
Biotech
SOM0226
Familial
amyloid
polyneuropathy
Phase
II
NCT02191826
2016
[
47
]
25
SOM3355
Huntington
chorea
Phase
II
NCT05475483
2024
[
48
]
27
SOM0061
COVID-19
Phase
II
e
2024
[
49
]
28
Neumora
BTRX-335140
Major
depressive
disorder
and
anhedonia
Phase
II
NCT04221230
2023
[
50
]
30
Xbiome
XBI-302
Acute-graft-versus-host
disease
Phase
I
NCT05352269
2022
[
51
]
31
Erasca
ERAS-007
Metastatic
colorectal
cancer
Phase
I/II
NCT05039177
2024
[
52
]
32
ERAS-601
Advanced
or
metastatic
solid
tumors
Phase
II
NCT04866134
2024
[
53
]
33
Nimbus
Therapeutics
NDI-034858
Psoriatic
arthritis
Phase
II
NCT05153148
2024
[
54
]
34
NDI
1150-101
Solid
tumor
Phase
I/II
NCT05128487
2024
[
55
]
35
Landos
Biopharma
BT-11
Crohn's
disease
Phase
II
NCT03870334
2023
[
56
]
36
NX-13
Ulcerative
colitis
Phase
I
NCT04862741
2023
[
57
]
37
C4X
discovery
INDV-2000
Orexin-1
Phase
I
NCT05694533
2024
[
58
]
38
Oncocross
OC514
Sarcopenia
PhaseI
NCT05264038
2023
[
59
]
39
Pharnext
PXT3003
CMT1A
Phase
II
NCT05092841
2025
[
60
]
40
Healx
HLX-0201
Fragile
X
syndrome
Phase
II
NCT04823052
2022
[
61
]
41
AbCellera
LY-CoV555
COVID-19
Completed
NCT05780268
2023
[
62
]
42
Insilico
Medicine
INS018-055
IPF
Phase
I
NCT05154240
2023
[
63
]
43
Schrodinger
SGR-1505
IPF
Phase
I
NCT05544019
2024
[
64
]
mCRPC: metastatic castration-resistant prostate cancer; SHP2: Src homology 2-containing protein tyrosine phosphatase 2; FGFR2:
fi
broblast growth factor receptor 2; PI3K
a
:
phosphatidylinositol
3-kinase
a
;
BOS:
bronchiolitis
obliterans
syndrome;
COVID-19:
coronavirus
disease;
CMT1A:
charcot-marie-tooth
disease,
type
1A;
IPF:
idiopathic
pulmonary
fi
brosis.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
3

prediction
model
using
the
DeepChem
framework.
Since
the
introduction
of
the
Lipinski
rule
of
fi
ve for drug
design,
predicting
the
early
ADMET
properties
of
lead
compounds
has
gained
increasing importance. Several studies have shown that training DL
models on large datasets of known compounds' ADMET properties
allows
for
the
automatic
identi
fi
cation
of
relationships
between
compound
characteristics
and
their
properties.
Well-trained
DL
models
can
then
predict
the
properties
of
novel
compounds,
thereby
accelerating
drug
discovery
and
development.
In
recent
years,
AI
has
begun
to
revolutionize
chemical
synthesis.
However,
the
lack
of
suitable
chemical
reaction
characterizations
and
the
scarcity of reaction data have limited the widespread application of
AI
in
reaction
prediction.
DL
can
address
these
challenges
by
automatically
identifying
and
extracting
patterns
from
chemical
synthesis
routes,
predicting
the
ef
fi
ciency
and
selectivity
of
new
synthetic routes, and accelerating the development of new drugs by
analyzing
large datasets
of
chemical
synthesis
reactions.
The
following
outlines
several
applications
of
DL
algorithms
in
virtual
screening
(VS):
convolutional
neural
networks
(CNNs)
are
particularly
effective
for
processing
image
data
such
as
molecular
structure
diagrams.
By
identifying
and
extracting
features
within
molecules,
such
as
atom
types,
positions,
and
chemical
bonds,
CNNs can predict the properties and activities of molecules. For VS
tasks involving sequence data (e.g., chemical molecular sequences),
recurrent
neural
networks
(RNNs)
are
well-suited.
RNNs
excel
at
capturing
long-term
dependencies
within
molecular
sequences,
improving
the
accuracy
of
property
predictions.
GANs
are
instru-
mental in generating novel molecular structures, a key advantage in
VS.
By
training
GANs,
it
is
possible
to
generate
molecules
with
desired properties, signi
fi
cantly reducing the need for experimental
validation.
Graph
neural
networks
(GNNs)
are
ideal
for
processing
graph-structured
data,
such
as
molecular
graphs
[
78
].
GNNs
can
model
the
relationships
between
atoms
and
chemical
bonds,
facilitating
more
accurate
predictions
of
molecular
properties.
For
tasks
involving
long
sequence
data,
such
as
multi-step
chemical
reaction
prediction,
Transformer
models
are
particularly
effective.
Transformers
can
capture
long-term
dependencies
within
se-
quences,
providing
enhanced
accuracy
in
predicting
molecular
properties
[
79
].
4.
AI
in
the
pharmaceutical
Drug
discovery
has
historically
relied
heavily
on
serendipity,
with
many
signi
fi
cant
breakthroughs
occurring
through
chance
observations
or
unintended
fi
ndings
[
80
].
However,
AI
offers
the
potential
to
remove
much
of
the
uncertainty
in
this
process,
dramatically
improving
the
chances
of
identifying
commercially
viable
drug
candidates
while
reducing both
costs
and
time.
A
pre-
dictive study in 2022 concluded that by heavily investing in AI, the
pharmaceutical industry could see a return on investment increase
of more than 45% [
81
]. The drug development process, which aims
to
identify
biologically
active
compounds
for
disease
treatment,
typically
begins
with
the
identi
fi
cation
of
molecular
targets,
fol-
lowed by the discovery of active drug candidates, and progresses to
the
optimization
of
lead
compounds
for
preclinical
and
clinical
trials, ultimately leading to regulatory approval. This process is not
only
time-consuming
but
also
high-risk
and
expensive,
with
the
average
cost
of
developing
a
new
drug
ranging
between
$100
million
and
$2
billion,
and
the
timeline
stretching
from
10
to
17
years.
Even
if
a
drug
candidate
successfully
passes
phase
I
clinical
trials,
it
has
only a
5%
chance of
reaching
the
market
[
82
].
Before
1980,
drug
discovery
primarily
relied
on
random
screening and empirical observations of natural product effects on
known
diseases
[
83
].
Although
inef
fi
cient,
this
method
led
to
the
discovery of several groundbreaking drugs, such as penicillin in the
1940s, which revolutionized
the treatment of previously incurable
diseases
like
tuberculosis
and
bacterial
infections.
Other
notable
discoveries
include
antihypertensive
drugs
like
prilosec,
lipid-
Fig.
2.
Overview
of
arti
fi
cial
intelligence
(AI),
machine
learning
(ML),
and
deep
learning
(DL).
CART:
classi
fi
cation
and
regression
tree;
KNN:
K-nearest
neighbor;
RBF:
radial basis
function;
SOM:
self-organizing
maps;
RF:
random
forest;
PCA:
principal
component
analysis;
LDA:
linear
discriminant
analysis;
DNN;
deep
neural
network;
FNNs:
feedforward
neural
networks;
GNNs:
graph
neural
networks;
CNNs:
convolutional
neural
networks;
RNNs:
recurrent
neural
networks.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
4

lowering statins, and anticoagulants like clopidogrel, which played
a
pivotal
role
in
managing
cardiovascular
diseases.
However,
after
the
1980s,
the
drug
discovery
process
improved
signi
fi
cantly
with
the
advent
of
high-throughput
screening
(HTS),
which
automates
the testing of thousands of compounds against molecular targets or
cellular assays. A notable milestone in HTS was the discovery of the
immunosuppressant
cyclosporine
A
in
1988
[
84
].
In
parallel,
re-
searchers
continued
to
invest
in
novel
methods
to
enhance
the
drug
discovery
process.
Computer-aided
drug
design
(CADD)
emerged as an effective approach to streamline the development of
new
drugs
[
85
].
CADD
utilizes
molecular
modeling
techniques
to
analyze the structure and interactions of numerous molecules, both
solid
and
non-solid,
against
pharmacological
targets,
assessing
their
activity,
toxicity,
and
bioavailability.
This
facilitates
better
planning and guidance throughout the drug discovery process [
86
]
(
Fig.
3
).
Several
signi
fi
cant
drugs,
including
the
anti-hypertensive
drug
captopril,
the
anti-human
immunode
fi
ciency
virus
(HIV)
drugs
saquinavir,
ritonavir,
and
indinavir,
and
the
protease
inhibi-
tor
boceprevir
for
hepatitis
C,
have
been
discovered
using
VS
techniques.
In
the
past
decade,
AI
has
gained
increasing
promi-
nence
in
CADD,
yielding
more
accurate
predictive
models.
The
emerging
fi
eld
of
AI
drug
design
(AIDD)
is
now
widely
recognized
within
the
pharmaceutical
industry.
It
is
anticipated
that
the
inte-
gration
of
AI
in
CADD
will
continue
to
revolutionize
drug
R
&
D,
making
future
drug
discovery
efforts
faster,
more
cost-effective,
and
more
successful
[
87
,
88
].
4.1.
Drug
characterisation
The encoding of molecules as
fi
xed-length strings or vectors is a
prerequisite
for
AI-driven
drug
molecule
research
[
89
].
Given
the
vast
chemical
space
of
drug
molecules,
selecting
appropriate
mo-
lecular features to accomplish speci
fi
c tasks is essential. Molecular
characterization,
also
referred
to
as
molecular
descriptors,
plays
a
pivotal
role
in
accurately
modeling
and
predicting
the
properties
and
biological
activity
of
small
molecules.
Such
characterizations
are
essential
for
applications
in
virtual
drug
screening,
compound
search,
ADME/T
prediction,
inverse
synthetic
route
planning,
and
other drug
discovery processes
[
90
,
91
].
The
article
discusses
various
types
of
molecular
fi
ngerprints
(MFPs),
such
as
simpli
fi
ed
molecular
input
line
entry
system
(SMILES),
substructure-based,
hash-based,
and
pharmacophore-
based
fi
ngerprints.
A
comparative
analysis
of
these
methods'
performance
metrics
(e.g.,
accuracy,
speed,
and
memory
usage)
on
benchmark
datasets
would
provide
valuable
guidance
for
selecting
the most suitable approach for speci
fi
c research needs.
Regarding GNNs for drug property prediction, the article brie
fl
y
highlights
their
potential.
Evaluating
the
performance
of
different
GNN architectures (e.g., graph convolutional network (GCN), graph
attention
network
(GAT),
and
Graph
SAmple
and
aggreGatE
(GraphSAGE))
on
benchmark
datasets,
alongside
an
assessment
of
their
interpretability
and
computational
ef
fi
ciency,
would
offer
critical insights into the advantages and limitations of each model.
4.1.1.
SMILES
SMILES is
widely employed
for drug
characterization,
encoding
molecular structures and geometric properties. Its linear molecular
representation
allows
SMILES
strings
to
be
processed
directly
as
text,
making
it
especially
useful
in
DL
models
for
various
drug
design tasks, such as inverse synthesis prediction via sequence-to-
sequence
(seq-2-seq)
methods.
Additionally,
SMILES'
ability
to
generate multiple representations of the same molecule by altering
atomic
order provides
an
advantage
for data
augmentation
[
92
].
4.1.2.
Molecular
fi
ngerprinting
A
MFP
is
a
bit
string
that
encodes
the
structural
or
pharmaco-
logical
properties
of
a
molecule
[
93
].
MFPs
are
widely
utilized
in
ligand-based similarity searches and quantitative structure-activity
relationship
(QSAR)
analysis,
especially
in
VS
for
drug
discovery.
Additionally,
DL-based
drug-target
interaction
(DTI)
prediction
models
frequently
use
MFPs
as
input
features
[
94
].
DTI
prediction
assists
researchers
in
assessing
the
ef
fi
cacy
and
safety
of
drug
candidates
early
in
the
process,
narrowing
the
search
for
thera-
peutic
targets,
and
expediting
the
discovery
and
development
of
new drugs
[
95
].
Drugs
and
targets
are
typically
represented
as
1D
sequences,
with DL models (such as CNNs, RNNs, and Transformers) employed
to
extract
features
and
make
predictions.
These
models
are
designed
to
effectively
represent
the
complex
structures
of
drugs
and targets, capturing their interactions. They also aim to enhance
the interpretability of predictions, provide insights into the model's
internal workings, and ensure the model generalizes well to unseen
data.
To
address
these
challenges,
new
DL
models
such
as
mutual
transformer-drug
target
af
fi
nity
(MT-DTA),
multi-scale
diffusion
and
interactive
learning-drug
target
af
fi
nity
(MDCT-DTA),
and
Transformer-graph
drug-target
af
fi
nity
prediction
(TGraphDTA)
Fig. 3.
Arti
fi
cial intelligence (AI) for the pharmaceutical industry. VS: virtual screening; ADMET: absorption, distribution, metabolism, excretion, and toxicity; PK: pharmacokinetics.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
5
have been proposed [
96
,
97
]. These models combine techniques like
diffusion
models,
graph
optimization,
and
interaction
learning
to
improve
feature
characterization
and
prediction
accuracy.
Furthermore,
the
interpretability
of
these
models
has
been
enhanced
by
visualizing
key
molecular
structures
upon which
the
models
focus.
Efforts
to
improve
feature
characterization
include
combining
the
3D
structures
of
proteins
with
drug
complex
structures,
as
well
as
developing
more
interpretable
models
using
explainable
AI
technologies.
To
improve
model
generalization,
larger-scale
biochemical
datasets
are
being
incorporated
[
98
].
Common
molecular
fi
ngerprinting
methods
include
substructure-
based,
hash-based,
and
pharmacophore-based
fi
ngerprints.
Notable
substructure-based
MFPs
include
the
molecular
access
system
(MAS)
and
PubChem
fi
ngerprints,
which
are
used
for
neighborhood
and
similarity
searches.
The
PubChemFP
encodes
881
structural
key
types
corresponding
to
the
substructures
of
all
compound
fragments
in
the
PubChem
database.
Hash-based
fi
n-
gerprints,
such
as
Daylight
FP,
Morgan
FP,
and
extended
connec-
tivity
fi
ngerprints (ECFPs), are widely used for compound similarity
analysis
[
99
].
Unlike
substructure-based
methods,
hash
fi
nger-
prints
convert
all
possible
fragments
into
values
using
a
hash
function.
Among
these,
ECFP,
a
recurrent
fi
ngerprint
based
on
Morgan's
algorithm,
is
commonly
used
as
input
for
DNNs
in
bioactivity
prediction
and
has
shown
strong
stability
[
100
].
Phar-
macophore
fi
ngerprints
assign
pharmacophore
types
to
atoms
in
the
chemical
structure,
generate
multiple
conformations,
and
construct
binary
fi
ngerprints
based
on
these
pharmacophores.
These
fi
ngerprints
are
used
as
descriptors
in
partial
least
squares
QSAR models. By capturing molecular features such as aromaticity,
hydrophobicity,
charge,
and
hydrogen
bond
donor/acceptor
prop-
erties,
pharmacophore
fi
ngerprints
enable
the
assessment
of
sim-
ilarity between target binding sites, considering energy-minimized
conformations of molecules to extract key pharmacophore features
[
101
].
4.1.3.
Molecular
characterisation
learning
The two methods of molecular characterization described above
yield
numerous
molecular
descriptors;
however,
the
results
are
often
constrained
by
the
domain-speci
fi
c
expertise
of
the
compu-
tational
chemist
and
the
choice
of
algorithm
employed
[
102
].
Determining
which
molecular
structures
and
properties
to
char-
acterize
for
optimal
downstream
processing
remains
a
challenge
[
103
].
GNNs
offer
a
comprehensive
and
generalized
approach
to
molecular
characterization.
By
representing
atoms
as
graph
nodes
and
chemical
bonds
as
graph
edges,
molecular
graphs
are
trans-
formed
from
abstract
mathematical
concepts
into
concrete
repre-
sentations
that
can
be
processed
by
computers.
These
graphs
are
mapped
onto
linear
data
structures,
such
as
matrices
or
arrays,
facilitating
computational
handling
[
104
].
In
a
GNN-based
molec-
ular graph, each atom and bond is associated with an initial feature
vector
within
a
feature
matrix.
The
atom's
feature
vector
typically
includes information about its local chemical environment, such as
atomic type, formal charge, and the number of attached hydrogens.
Bond
features
may
include
the
adjacency
matrix,
bond
type,
shortest path
length,
and
the
presence
or absence
of
speci
fi
c
rings
[
105
].
GNNs
can
automatically
learn
task-speci
fi
c
molecular
rep-
resentations
through
graph
convolution,
eliminating
the
need
for
traditional
manual
descriptors
or
fi
ngerprints,
and
they
demon-
strate
high
accuracy
in
predicting
compound
properties
[
106
].
Predicting
a
molecule's
chemical
properties
or
biological
activity
directly
from
its
structure
has
long
been
a
focus
of
interest
within
the
chemical
community
[
107
].
The
GNN-based
fi
ngerprinting
method,
Neural
FP,
utilizes
graph
CNNs
to
learn
molecular
repre-
sentations
directly
from
molecular
graphs.
The
Weave model
con-
siders both atoms and chemical bonds within the molecular graph,
optimizing
atomic
and
atomic
pair
features,
and
has
proven
effec-
tive
in
predicting
water
solubility,
biological
activity,
and
toxicity.
Similarly,
Attentive
FP
introduces
a
graph
attention
mechanism
to
model
node
information,
capturing
local
and
non-local
features
of
chemical
structures,
such
as
intramolecular
hydrogen
bonds
and
aromatic
systems.
This
enables
Attentive
FP
to
excel
at
learning
molecular
representations
for
a
wide
range
of
properties
[
108
].
Beyond
drug
property
prediction
and
characterization,
GNNs
are
also
applicable
in
areas
such
as
ab
initio
drug
design,
interaction
prediction,
and
inverse
drug
discovery [
109
].
4.2.
Target
discovery
and
validation
Targeted
drug
discovery
remains
a
cornerstone
of
pharmaceu-
tical
development.
When
a
drug's
target is
known, designing
drug
screening
experiments
to
identify
therapeutics
acting
on
that
protein
target
becomes
more
straightforward
[
110
].
However,
failing
to
identify
a
target
accurately
can
result
in
signi
fi
cant
R
&
D
investment
losses.
For
instance,
clinical
trials
by P
fi
zer,
Roche,
and
Merck Sharp
&
Dohme for cholesteryl ester transfer protein (CETP)
inhibitors,
a
lipid-lowering
target,
ended
in
failure
[
111
].
Conversely,
the
discovery
of
programmed
cell
death
1
(PD-1)
has
revolutionized
biomolecule
and
tumor
immunotherapy,
and
it
is
projected
that
more
than
20
PD-1
products
will
reach
the
market
globally within the next 2
3 years. On the other hand, even when a
novel druggable protein target is discovered, the path to bringing a
new
chemical
entity
to
market
is
fraught
with
signi
fi
cant
chal-
lenges,
especially
concerning
development
time
and
cost.
Identi-
fying new targets or indications for an existing drug, however, can
substantially reduce development expenses [
112
]. One of the most
notable
examples
of
drug
repurposing
is
sildena
fi
l,
and
AI
tech-
nology
has
the
potential
to
transform
such
serendipitous
discov-
eries
into
more
systematic
successes
[
113
].
By
combining
systems
biology with AI algorithms, correlations between multi-omics data
and patient clinical health information can be mined. Additionally,
using NLP to retrieve and analyze unstructured data from literature,
patents,
and
clinical
reports
can
help
uncover
potential
disease-
relevant
pathways,
proteins,
and
mechanisms.
This
approach
aids
in the identi
fi
cation of new targets for drug development, whether
for
novel chemical
entities
or
repurposed
drugs
[
114
,
115
].
i) Comparison of reverse docking software: the article mentions
reverse
docking
as
a
method
for
target
discovery.
A
performance
comparison
of
various
reverse
docking
software
(e.g.,
AutoDock,
Glide, and Rosetta) on a benchmark dataset, considering factors like
ease of use and scalability, could help researchers identify the most
suitable
tool for
targeting speci
fi
c
proteins.
ii)
Comparison
of
protein
structure
prediction
methods:
the
article
discusses
AlphaFold
as
a
protein
structure
prediction
method.
Comparing
its
performance
(e.g.,
Global
Distance
Test
(GDT)
score) with
other
methods such as
I-TASSER or
Rosetta
on a
benchmark
dataset
could
provide
valuable
insights
into
the
strengths
and
limitations
of
different
protein
structure
prediction
approaches.
4.2.1.
Systems
biology
approach
By examining the
interrelationships and
interactions
of various
components within biological systems at the molecular level, such
as
gene
and
protein
networks
related
to
cell
signaling,
metabolic
pathways,
organelles,
cells,
physiological
systems,
and
organisms,
systems
biology
aims
to
create
comprehensive
models
and
com-
plete
organism
maps
[
116
].
Network-based
approaches
infer
new
protein
phenotypes
or
associations
by
linking
proteins/genes
to
different
network
pathways
[
117
].
However,
the
intricate
complexity
of
biological
network
interactions
presents
challenges
in
constructing
network-based
models
for
disease
classi
fi
cation,
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
6
personalized
medicine,
and
prognosis.
These
models
often
fail
to
provide
stable
pathway
signatures
for
speci
fi
c
phenotypes
or
reli-
able
biomarkers
of
disease,
hindering
the
creation
of
unbiased,
data-driven
networks
for
identifying
biomarkers,
targets,
and
dis-
eases
[
118
].
Utilizing
Bayesian
AI
analysis
to
integrate
molecular
pro
fi
les
from
multi-omics
data
(such
as
genomics,
proteomics,
lipidomics,
and
metabolomics)
with
clinical
health
information
enables
the
construction
of
causal
inference
networks.
By
comparing
the
dif-
ferences
between
"health"
and
"disease"
network
graphs,
disease
drivers
(targets
and
biomarkers)
can
be
identi
fi
ed.
This
approach
led
to
the
discovery
of
the
novel
tumor
target
BPM42522,
its
lead
molecule,
and
its
anticancer
mechanism
of
action
[
119
,
120
].
The integration of knowledge mapping techniques with systems
biology
to
build
biomedical
knowledge
graphs
has
increasingly
become pivotal in medical practice and research. These graphs help
to simplify complex biological systems and pathological processes,
offering
a
clearer
understanding
of
underlying
principles.
When
combined
with
disease-speci
fi
c
contexts,
biomedical
knowledge
graphs
facilitate
drug
repurposing
and
mechanistic
analysis
of
emerging
human
diseases,
such
as
coronavirus
disease
2019
(COVID-19)
[
121
].
BenevolentAI
has
introduced
a
judgment-
enhanced
cognitive
system
(JECS)
that
uses
AI
tools
and
biomed-
ical
knowledge
graphs
to
identify
potential
drug
candidates.
By
discovering
new
connections
between
vast
amounts
of
unstruc-
tured data, such as disease, drug, and clinical trial information, JECS
enables
drug
redirection
and
assists
scientists
in
identifying
new
indications
for
existing
drugs
[
122
].
Similarly,
MindRankAI
has
developed
PharmKG39,
a
multi-relational
biomedical
knowledge
graph of
drug-disease
associations,
which
integrates
over
500,000
relationships
between
genes,
drugs,
and
diseases.
Using
a
hetero-
geneous graph attention neural network, PharmKG39 incorporates
29 relationship categories and over 8000 ambiguous entities, each
enriched
with
domain-speci
fi
c
information
from
datasets
on
gene
expression,
chemical
structure,
and
disease
word
embedding,
preserving
both
semantic
and
biomedical
features
[
123
].
The
development
of
the
in
silico
pathway
activation
network
decomposition
analysis
(iPANDA)
method
by
scientists
at
Insilico
Medicine
represents
a
signi
fi
cant
advancement
in
pathway
acti-
vation
analysis
[
124
].
iPANDA
is
designed
to
extract
biologically
relevant
features
from
large-scale
transcriptomic
and
proteomic
data, offering a powerful approach for biomarker identi
fi
cation. The
method uses
gene expression
data to assess fold changes between
tumor
samples
and
normal
group
averages,
incorporating
gene
importance
factors
to
quantify
the
in
fl
uence
of
genes
on
speci
fi
c
pathways.
However,
the
measure
of
gene
centrality
varies
across
algorithms, and this variation can lead to highly inconsistent results
[
125
].
In
this
particular
study,
iPANDA
integrates
the
degree
of
differential gene expression with pathway topology decomposition
into a uni
fi
ed network model. By utilizing statistical and topological
weights,
gene
importance
is
estimated.
Additionally,
gene
co-
expression
modules
are
introduced,
and
the
topological
co-
ef
fi
cients
of
these
modules
are
computed
to
obtain
co-expression
data.
This
data
is
then
combined
with
the
gene
importance
fac-
tors
to
calculate
pathway
activation
scores.
Leveraging
iPANDA,
Insilico
Medicine
has
developed
a
new
target
discovery
platform
called PandaOmics, which has been instrumental in identifying and
prioritizing
over
20
novel
targets
for
IPF.
The
platform
compares
histology
data
from
patients
with
fi
brosis
to
healthy
individuals,
identifying signi
fi
cant differences and utilizing iPANDA technology
to
pinpoint
pathways
that
may
affect
these
differences.
Following
this,
target
safety
and
future
value
were
assessed
through
target
knockout data, leading to the identi
fi
cation of promising targets for
IPF
treatment.
The
project
has
progressed
to
clinical
stages,
marking a signi
fi
cant milestone in the development of therapies for
IPF
[
126
].
4.2.2.
Target
structure-based
approaches
Target
con
fi
rmation
(i.e.,
target
selection
or
prioritization)
in
drug
development
remains
an
uncertain
process,
necessitating
precise mapping of interactions between approved drugs and their
ef
fi
cacy targets, i.e., the speci
fi
c proteins or molecules on which the
drugs
exert
therapeutic
effects
[
127
].
Structure-based
computa-
tional methods for target discovery serve as valuable complements
to
experimental
strategies,
such
as
reverse
docking,
pharmaco-
phore
modeling,
binding
site
similarity,
and
fi
ngerprint-based
in-
teractions
[
128
].
Among
these,
reverse
docking
has
proven
particularly
effective,
not
only
for
target
validation
but
also
for
predicting toxicity and side effects, as well as for uncovering novel,
previously
unidenti
fi
ed
targets
for
drugs
or
natural
compounds
[
129
]. The Potential Drug Target Database (PDTD), a comprehensive
database
for
reverse
docking,
has
been
used
to
identify
potential
targets
for
compounds
such
as
tea
polyphenols
and
ginsenosides.
However,
the
reverse
docking
method
is
constrained
by
the
avail-
able target structure dataset. The PDTD, released in 2008, includes
approximately
1100
protein
entries
with
3D
structures,
sourced
from
literature
and
various
online
repositories
(e.g.,
Therapeutic
Target Database (TTD), DrugBank, and Thomson Pharma), covering
830
known
or
potential
drug
targets
[
130
].
Notably,
only
approxi-
mately 11% of the human proteome has been annotated with small
molecule
probes,
leaving
a
signi
fi
cant
portion
of
proteins,
approx-
imately one-third
of
the
human proteome,
still
uncharacterized
in
terms of
their
biological
functions
and
roles in
disease.
Three
methods,
nuclear
magnetic
resonance
(NMR),
X-ray
crys-
tallography, and cryo-electron microscopy, are widely employed for
protein
structural
resolution
and
have
yielded
signi
fi
cant
insights
into
protein
structures
and
drug
receptors.
Many
drugs,
such
as
angiotensin-converting enzyme
(ACE)
inhibitors, have
entered
clin-
ical practice based on structural data derived from these techniques
[
131
e
133
].
However,
these
methods
are
not
without
their
chal-
lenges.
Cryo-electron microscopy equipment
is priced between
$20
million and $60 million, while the synchrotron light sources required
for X-ray crystallography can cost several hundred million dollars to
construct
[
134
].
Additionally,
the
time
required
to
resolve
a
protein
structure
can
range
from
weeks
or
months
to
several
years,
in
fl
u-
enced by factors such as sample availability and protein complexity
[
135
].
In
light
of
these
constraints,
AI-based
protein
structure
pre-
diction
algorithms
present
a
promising
complement
to
traditional
protein
information-driven
target
validation
approaches.
AlphaFold
2, for instance, achieved a median GDTscore of 92.4 across all targets,
closely
approximating
the
quality
of
results
from
gold-standard
experimental
methods
like
X-ray
crystallography
[
136
].
AlphaFold
3, developed
by
DeepMind, employs
a
DNN architecture
trained
on
170,000
protein
structures
from
the
Protein
Data
Bank
(PDB)
to
predict inter-amino acid distances and torsion angles between bonds
in
protein
structures
[
137
].
The
methods
and
architecture
behind
AlphaFold
3
were
recently
published, and
in collaboration with
Eu-
ropean
Molecular
Biology
Laboratory-European
Bioinformatics
Institute (EMBL-EBI), AlphaFold 3's predictions, covering 98.5% of the
human
proteome,
have
been
made
publicly
available
for
the
scien-
ti
fi
c
community
[
138
].
Although
AlphaFold
3
still
struggles
with
accurately
modeling
side-chain
structures
and
dynamics
and
faces
challenges
in
predicting
the
structures
of
multi-domain
proteins,
protein
complexes,
and
membrane
proteins,
its
extensive
protein
structure
library
provides
a
nearly
comprehensive
reverse
docking
dataset.
This
resource
enables
the
identi
fi
cation
of
potential
target
proteins,
thus
helping
to
mitigate
the
limitations
of
existing
target
structure
datasets
in
reverse
docking.
As
a
result,
reverse
docking,
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
7
bolstered
by
AlphaFold
3's
insights,
has
the
potential
to
become
a
truly
invaluable
tool
in
drug
discovery,
advancing
the
fi
eld
signi
fi
-
cantly [
139
e
141
].
4.3.
Small
molecule
drug
discovery
In
recent
years,
DL
has
signi
fi
cantly
impacted
fi
elds
such
as
image
analysis
and
NLP.
Motivated
by
these
advancements,
computational
chemists
are
increasingly
employing
generative
models to design new molecules and predict their properties [
142
].
The
vast
chemical
space
encompasses
approximately 10
60
to 10
100
possible
small
molecules,
making
drug
discovery akin
to
fi
nding
a
needle in a haystack, as researchers must identify compounds that
satisfy
multiple
criteria,
such
as
biological
activity,
metabolic
sta-
bility,
and
potency.
Consequently,
only a
tiny
fraction
of
this
theo-
retical
chemical
space
can
be
explored
through
traditional
experimental
approaches.
Computer
modeling
techniques,
how-
ever,
are
proving
invaluable
in
enhancing
the
biological
screening
of
large
compound
libraries
and
optimizing
synthetic
routes
to
complementary
compounds,
thus
playing
a
pivotal
role
in
early
drug
discovery [
143
].
i)
Comparison
of
molecular
generator
techniques:
the
article
discusses the use of GANs for molecule generation. Comparing the
performance
of
various
GAN
architectures,
such
as
conditional
generative
adversarial
network
(CGAN)
and
style-based
generator
architecture
for
generative
adversarial
networks
(StyleGAN),
on
benchmark datasets,
and evaluating their ability to generate novel
and
biologically
active
molecules,
could
assist
researchers
in
selecting the most effective technique for drug discovery purposes.
ii) Comparison of synthetic route planning methods: the article
also
addresses
template-based
and
template-free
approaches
for
synthetic
route
planning.
A
comparison
of
their
performance,
considering
factors
such
as
ef
fi
ciency,
selectivity,
and
by-product
formation,
on
benchmark
datasets,
could
guide
researchers
in
choosing
the
most
suitable
strategy
for
synthesizing
their
target
molecules.
4.3.1.
Molecular
generator
techniques
A key aspect of compound design and predictive modeling is the
selection
of
appropriate
molecular
representations.
Text
or
string
encoding
of
molecules
is
computationally
inexpensive
and
commonly
used
in
molecular
generators.
In
generative
modeling,
SMILES-based
string encoding
typically generates a
token
for each
atom,
which
is
then
converted
into
a
“
one-hot
”
string
representa-
tion [
144
]. A
generative model using one-hot encoding produces a
distribution
for
each
token,
which
is
sampled
to
generate
a
new
structure
in
SMILES
format.
Alternatively,
graph-based
generative
modeling,
using techniques
like GCNs or DL to generate molecular
structures,
is
an
emerging
area
of
research.
Rule-based
graph
generative models often yield structurally correct molecules but are
computationally
intensive.
The
integration
of
fl
exible
neural
network
architectures
and
diverse
molecular
representations
has
led
to
the
development
of
various
innovative
approaches
for
mo-
lecular
generation
[
145
].
4.3.2.
Synthetic
route
planning
Computer-aided
synthetic
planning
(CASP)
has
its
roots
in
the
pioneering
work
of
E.J.
Corey,
who
formalized
the
concept
of
"in-
verse synthetic analysis" in the late 1960s [
146
]. CASP incorporates
the principles of inverse synthetic analysis to help synthetic organic
chemists
identify
the
most
ef
fi
cient
and
cost-effective
synthetic
routes,
predict
selectivity
and
by-products,
and
suggest
optimal
reaction
conditions.
Over
the
decades,
computational
methods
have
evolved
from
expert
systems
based
on
manually coded
reac-
tion
rules
and
templates
to
data-driven,
AI-assisted
synthesis
planning
[
147
].
Modern
AI
algorithms
are
now
available
to
recommend feasible synthetic routes for a wide range of reactions,
with
or
without
reaction
templates,
operating
at
either
the
mech-
anistic
or
global
reaction
level.
These
methods
utilize
molecular
representations
such
as
fi
ngerprints,
graphs,
or
even
SMILES
strings.
CASP
is
instrumental
in
enabling
chemists
to
make
better
decisions,
thereby
increasing
ef
fi
ciency and
productivity,
reducing
synthetic
failures,
and
accelerating
the
design-make-test-assess
(DMTA)
cycle
in
drug
discovery [
148
].
Rule-based
approaches
rely
on
expertly coded
rules
and
heuris-
tics
extracted
from
reaction
databases
and
literature
to
suggest
synthetic
routes,
often
referred
to
as
"template
methods".
In
such
approaches, reaction rules are manually curated, which is limited by
the
inability
to scale with
the
exponential
growth of chemical
liter-
ature
and
by
the
fi
nite
knowledge
base
that
cannot
be
fully
comprehensive. Synthia (Chematica) is an inverse synthesis software
that leverages a library of expertly coded rules for chemical synthesis
planning [
149
]. To address the limitations of the rule-based system,
Synthia
incorporates
computational
methods
to
automate
the
extraction
of
reaction
rules
from
reaction
datasets.
Its
template
extraction
algorithm,
based
on
Ambit-SMIRKS,
is
speci
fi
cally
designed
to
describe
chemical
reactions
and
has
accumulated
an
expert-coded rule base of approximately 50,000 rules over 15 years
[
150
]. Synthia's core algorithm utilizes a decision tree where various
conditions
determine
the
range
of
possible
substituents
or
atom
types.
A
scoring
function
and
dynamic
planning
algorithm
then
construct complete synthetic pathways by making decisions for each
inverse synthetic step, enabling the proposal of synthetic routes for
all targets within 15
e
20 min. In 2024, L
opez-Ch
avez et al. [
151
] used
Synthia
to
design
synthetic
pathways
for
eight
structurally
diverse
and synthetically challenging molecules, marking the
fi
rst successful
use
of
synthetic
planning
software
to
guide
multi-step
synthetic
routes.
The
highest-scoring
synthetic
route
was
selected
to
synthe-
size the targets, achieving yields of up to 98%. Notably, the synthetic
route
proposed
by
Synthia
signi
fi
cantly
differed
from
the
original
patent-disclosed route, providing higher yields with fewer synthetic
steps
[
152
].
In
recent
years,
AI
techniques
have
been
applied
to
the
extraction
of
reaction
rules.
Segler
et
al.
[
153
]
pioneered
a
neural-
symbolic
approach
to
autonomously extract
inverse
synthesis
rules
from
the
Reaxys
database
without
expert
input.
These
rules
were
then integrated with modern Monte Carlo tree search algorithms for
reaction prediction to identify the most promising inverse synthesis
routes.
However,
template-based
approaches
present
challenges
such
as
high
computational
costs
and
incomplete
rule
coverage,
which restrict their scalability [
154
,
155
].
To
overcome
the
limitations
of
the
template-based
approach,
the
template-free
approach
draws
from
NLP
and
frames
synthetic
prediction,
whether
forward
or
inverse,
as
a
Seq-2-Seq
mapping
problem
[
156
].
Since
molecules
can
be
represented
as
SMILES
strings,
each
chemical
reaction
can
be
encoded
as
a
sentence,
thereby
treating
it
as
a
chemical
language
translation
issue
[
157
].
The
fi
rst
template-free
approach
to
inverse
synthesis
analysis
employed a Seq-2-Seq model, fully data-driven and trained end-to-
end
on
a
subset
of
experimental
reactions
with
labeled
reaction
types.
This
model
consists
of
a
bidirectional
long
short-term
memory
(LSTM)
encoder-decoder,
augmented
with
an
attention
mechanism
that
maps
the
SMILES
representation
of
reactants
to
those
of
the
products
[
158
,
159
].
The
performance
of
this
method
has
been
shown
to
be
comparable
to that
of
a
baseline
rule-based
expert
system.
4.4.
Small
molecule
drug
design
and
optimization
The
review
discusses
various
scoring
functions
used
in
structure-based
VS
(SBVS)
and
suggests
that
comparing
their
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
8
performance
metrics,
such
as
enrichment
rate
and
selectivity,
on
benchmark
datasets
could
assist
researchers
in
selecting
the
most
appropriate
scoring
function
for
their
target
proteins.
Similarly,
it
highlights QSAR and pharmacophore-based approaches for ligand-
based
VS
(LBVS),
with
performance
metrics
like
accuracy and
ef
fi
-
ciency
serving
as
key
criteria
for
identifying
the
best
method
for
target
molecules.
Additionally,
the
article
touches
on
different
ML
and
DL
techniques
for
ADMET
prediction,
proposing
that
perfor-
mance
metrics
such
as
accuracy
and
robustness
could
guide
re-
searchers
in
choosing
the
most
suitable
model
for
their
compounds.
4.4.1.
SBVS
SBVS,
also
referred
to
as
target-based
VS
(TBVS),
is
a
powerful
and
promising
CADD
method.
This
approach
predicts
the
interac-
tion
between
a
target
protein
and
a
vast
library
of
compounds
by
utilizing the 3D structure of the target. Compounds are scored and
ranked
based
on
their
predicted
af
fi
nity
for
the
target's
receptor
binding
site,
facilitating
the
identi
fi
cation
of
those
most
likely
to
exhibit pharmacological activity against the molecular target [
160
].
Molecular
docking,
a
central
technique
in
SBVS,
examines
the
geometric compatibility between ligands and their targets. Docking
became
especially
valued
for
its
low computational
cost,
ability
to
conduct
virtual
testing
before
molecular
synthesis,
and
its
effec-
tiveness
in
time
and
cost-saving.
However,
while
SBVS
is
widely
used, its effectiveness can be limited by system-speci
fi
c challenges.
The complexity of ligand-receptor binding interactions complicates
accurate
parameterization,
leading
to
dif
fi
culties
in
predicting
binding sites and classifying compounds. This often results in high
false-positive
and
false-negative
rates
[
161
].
Accurate
SBVS
de-
pends heavily on two main components: the search algorithm and
the
scoring
function.
The
search
algorithm
systematically explores
ligand orientations and conformations at the binding site, while the
scoring
function
predicts
the
binding
af
fi
nity
between
the
target
and
its
candidate ligands.
Both
components
are
critical
to the
suc-
cess
of
docking
protocols
[
162
].
Scoring
functions
play a
critical
role
in
molecular
docking
with
three
primary
applications:
i)
determining
the
binding/alteration
sites of targets and ligands, as well as the binding conformation; ii)
predicting
the
binding
af
fi
nity
between
proteins
and
ligands;
and
iii)
optimizing
potential
ligands
[
163
].
Traditionally,
scoring
func-
tions
are
categorized
into
three
main
types:
force
fi
eld-based,
empirical,
and
knowledge-based
scoring
functions.
In
recent
years,
ML-based
scoring
functions
have
emerged
as
a
fourth
type
[
164
].
While
traditional
scoring
functions
are
widely
used,
they
have
notable
limitations,
such
as
inadequate
consideration
of
conformational entropy (the
fl
exibility of the protein) and solvation
energy
[
165
].
With
the
abundance
of
experimental
data
available,
AI
algorithms
can
now
build
non-prede
fi
ned,
data-driven
scoring
functions.
These
models
implicitly
learn
the
eigenvectors
of
protein-ligand
binding
and
their
non-linear
relationships
with
af-
fi
nity.
Several
researchers
have
successfully
employed
ML-based
scoring
functions
to
enhance
SBVS
algorithms.
Notable
examples
include
random
forest
(RF)-Score-VS
and
SFCscore
RF
based
on
RF,
support
vector
regression-knowledge-based/-physico-chemical
properties
(SVR-KB/-EP)-score
and
ID-score
based
on
support
vector machine (SVM), and NNScore 2.0 and CScore based on early
arti
fi
cial
neural
networks
[
166
e
170
].
While
traditional
ML
approaches
still
depend
on
expert
knowledge and feature engineering, the rise of DL algorithms offers
a promising direction for scoring function modeling [
171
]. CNNs, for
instance, can automatically extract features directly from 2D or 3D
structures
to
predict
the
binding
af
fi
nity
between
proteins
and
ligands.
The
3D
lattices
of
protein-ligand
structures
generated
by
docking
can
serve
as
input
to
CNN
models.
From
these
lattices,
relevant
features,
such
as
complex
atom
types,
partial
atomic
charges,
and
interatomic
distances,
are
automatically
learned
and
extracted. These features are then used to build regression models
for
predicting
af
fi
nity
or
classi
fi
cation
models
for
predicting
bind-
ing
or
non-binding
interactions.
CNN-based
models
have
shown
better
predictive
performance
than
traditional
docking
methods
[
172
].
The
introduction
of
DL
techniques,
particularly
CNNs,
has
revitalized
SBVS.
Traditional
scoring
function
approaches
use
pre-
de
fi
ned
theories
to
design
functions
based
on
linear
relationships.
In
contrast,
AI
techniques
can
implicitly
capture
intermolecular
binding interactions that are challenging to model explicitly. While
DL-generated
scoring
functions
may
not
always
outperform
established ML methods, further optimization of training ef
fi
ciency
and interpretability is underway. Despite this, the incorporation of
DL has already led to signi
fi
cant improvements in existing docking
tools.
As
a
result,
SBVS,
powered
by
AI
and
DL,
is
expected
to
become
one
of
the
most
promising
techniques
in
the
drug
discov-
ery
process
in
the
near
future [
173
].
4.4.2.
LBVS
LBVS
operates
on
the
premise
that
structurally
similar
com-
pounds
exhibit
comparable
biological
activities.
Commonly
employed
methods
in
LBVS
include
QSAR,
pharmacophore
modeling,
and
structural
similarity
matching
[
174
].
The
QSAR
model,
developed
over
the
past
half-century,
establishes
a
mathe-
matical
correlation
between
a
compound's
molecular
properties
(e.g.,
polarity,
lipophilicity,
electrical
and
spatial
characteristics,
or
speci
fi
c
structural
features)
and
its
biological
activity
indicators
(such as receptor af
fi
nity, inhibition constants, or rate constants). A
re
fi
nement
of
QSAR,
3D-QSAR,
directly
derives
binding
af
fi
nities
from
the
3D
structure
of
the
compound,
while
comparative
mo-
lecular
fi
eld
analysis
(CoMFA)
is
a
pivotal
method
for
3D
confor-
mational
analysis.
The
3D
pharmacophore
model
involves
the
conformational analysis and molecular stacking of a series of active
compounds to identify critical moieties that in
fl
uence their activity.
In
pharmacophore-based
VS,
3D
pharmacophores,
derived
from
active
ligands,
target-ligand
complexes,
or
protein
structures,
are
screened
against
a
virtual
compound
library,
retrieving
molecules
that
satisfy
the
pharmacophore's
criteria
[
175
e
177
].
Structural
similarity matching identi
fi
es compounds with analogous activities
or mechanisms by comparing molecular descriptors or
fi
ngerprints.
AI-based
VS
models
leverage
molecular
descriptors
derived
from physicochemical properties and/or topological
fi
ngerprints to
build
regression
or
classi
fi
cation
models
of
activity.
This
approach
offers
greater
fl
exibility
in
LBVS,
eliminating
dependency
on
program-speci
fi
c functionalities. Bayesian algorithms, SVM, RF, and
arti
fi
cial
neural
networks
have
been
extensively
used
to
construct
QSAR
models,
driving
numerous
successful
applications
in
LBVS
[
178
].
DNNs
have
outperformed
traditional
ML
methods
like
Bayesian, RF, and SVM in predictive accuracy. Multi-task DNNs have
further
enhanced
performance,
with
applications
on
200
distinct
targets
for
large-scale
screening.
The
DeepTox
method
[
179
],
a
multi-task DNN-based toxicity prediction model, triumphed in the
2014
Tox21
dataset
challenge,
which
involved
predicting
com-
pound toxicity using a dataset of 12 high-throughput assays across
12,000
compounds
[
180
].
The
introduction
of
GNNs
in
molecular
prediction
allows
for
a
more
comprehensive
and
generalized
rep-
resentation
of
molecules,
aiding
in
the
automatic
extraction
of
relevant molecular features for predictive modeling [
181
]. Recently,
molecular prediction models for antimicrobial activity based on the
message
passing
neural
network
(MPNN)
identi
fi
ed
eight
novel
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
9
antimicrobial
molecules
from
a
database
of
over
107
million
com-
pounds,
including
the
discovery
of
halicin,
a
new
antibiotic
that
inhibits
Escherichia
coli
(
E.
coli
)
growth
[
182
].
The integration of pharmacophore concepts with AI techniques
is
still
evolving,
with
future
research
likely
focusing
on
using
pharmacophore features as molecular descriptors for AI models or
employing AI methods to generate pharmacophores from extensive
datasets.
For
example,
when
Pharm-IF,
a
pharmacophore-based
interaction
fi
ngerprint,
was
used
as
input
to
several
ML
algo-
rithms
for
ranking
small
molecule
docking
poses,
the
SVM-based
model
outperformed
other
algorithms
and
docking
scores
in
terms of enrichment rates [
183
]. It is anticipated that AI algorithms,
in conjunction with increasingly re
fi
ned molecular characterization
methods,
will
soon
dominate LBVS techniques.
4.4.3.
Predicting
pharmacogenicity
ADMET
properties
are
essential
indicators
in
determining
whether
a
small
molecule
can
develop
into
a
viable
drug,
addressing
key
pharmacokinetic
and
toxicological
concerns,
such
as the drug's ability to be effectively absorbed and reach the target
tissue.
Many
clinical
trial
failures
are
attributed
to
inadequate
ADMET
pro
fi
les
in
drug
candidates.
Early-stage
ADMET evaluation
can
effectively
mitigate
safety
and
ef
fi
cacy
concerns,
thereby
increasing
the
success
rate
of
drug
development
[
184
].
However,
traditional experimental methods for ADMET evaluation are costly
and
time-consuming,
limiting
early
insights
into
compounds
and
delaying
further
biological
validation.
The
advancement
of
computational technology, cheminformatics, and the accumulation
of
experimental
drug
data
have
facilitated
the
development
of
ADMET prediction models using ML and DL. These models can learn
the
relationship
between
chemical
structures
and
pharmacoki-
netics
from
ADMET
data,
helping
medicinal
chemists
avoid
exploring
suboptimal
chemical
spaces
and
enabling
the
identi
fi
-
cation
of
promising
drug
candidates
[
185
].
ADMET prediction is a critical component of drug discovery and
development.
Leading
global
institutions
and
companies
are
increasingly
integrating
traditional
wet
lab
experiments
with
computational
methods
to
aid
in
the
analysis
of
ADMET
pro
fi
les,
resulting in the development of numerous computer-aided ADMET
software,
databases,
and
online
services
[
184
].
For
instance,
the
QikProp
module
from
Schr
€
odinger
software
predicts
key
pharma-
cokinetic
parameters
such
as
logP,
logS,
Caco-2
cell
permeability,
serum
protein
binding,
and
human
ether-a-go-go
related
gene
(hERG)-K
ion
channel
blocking.
GastroPlus,
widely
adopted
by
regulatory
authorities
like
the
U.S.
FDA
and
the
China
National
Medical
Products
Administration
(NMPA),
forecasts
pharmacoki-
netic parameters including physicochemical properties, absorption,
distribution,
metabolism,
and
drug
behavior
after
ocular
and
pul-
monary administration [
186
]. The SwissADME molecular modeling
team
at
the
Swiss
Bioinformatics
Institute
calculates
physico-
chemical descriptors to predict pharmacokinetic properties such as
oral
bioavailability,
blood-brain
barrier
permeability,
and
in-
teractions with metabolic enzymes, in addition to offering a suite of
widely used online tools for ADMET prediction [
187
]. ADMETlab 3.0
utilizes
MFPs
like
MAS
and
ECFPs
to
train
ML
models
such
as
RF,
SVM,
and
naive
Bayes
for
the
classi
fi
cation
and
regression
predic-
tion
of
various
ADMET
properties.
Additionally,
ADMET
SAR
em-
ploys
MACCS
fi
ngerprints
to train
SVM models,
achieving
superior
prediction performance in 22 classi
fi
cation tasks, with the adoption
by DrugBank, a prominent drug database. Although ML-based tools
are
the
most
widely
used,
employing
MFPs
and
descriptors
as
features
can
lead
to
signi
fi
cant
loss
of
molecular
structural
infor-
mation, potentially limiting the predictive accuracy of these models
[
188
,
189
].
DL-based
ADMET
prediction
methods
are
capable
of
autono-
mously extracting feature representations from input data to model
more
intricate
relationships.
As
demonstrated
in
the
2020
Kaggle
Competition,
DNNs
outperformed
RF
models
by an
average
of
10%
across
15
large
analytical
datasets
[
190
].
Researchers
from
leading
pharmaceutical
companies
such
as
Vertex
Inc.,
Eli
Lilly
&
Co.,
and
Bayer
AG
have
also
reported
that
DNNs
either
match
or
slightly
surpass traditional ML models when trained on proprietary ADMET
datasets
[
191
e
194
].
The recent rise
of
GNNs
has
introduced
a
new
paradigm
in
ADMET
model
design.
GNNs
represent
molecules
as
graph
structures
and,
through
data-driven
training,
convert
mo-
lecular
structural
information
into
low-dimensional
continuous
vectors,
offering
a
more
informative
and
compact
representation
compared
to
traditional
high-dimensional
sparse
MFPs
[
195
,
196
].
The ef
fi
cacy of GNN models in predicting drug properties has been
validated
by
frameworks
such
as
Molecule-Net
and
Chemi-Net.
Chemi-Net,
a
fully
data-driven,
domain-knowledge-free
deep
GCN
developed
in
collaboration
with
Amgen,
outperformed
Amgen's Cubist ML program across 13 datasets in large-scale ADME
property prediction, demonstrating its superior predictive accuracy
and potential to accelerate drug discovery [
197
]. Additionally, GNNs
can
leverage
interpretable
methods,
such
as
self-attention-based
message-passing
neural
network
(SAMPN),
a
message-passing
neural
network
based
on
a
self-attention
mechanism.
SAMPN
has
been shown to outperform both conventional GNNs and RF models
in predicting lipophilicity and water solubility, while also enabling
the
visualization
of
atomic
contributions
to
the
predicted
proper-
ties
through
attention
coef
fi
cients
[
198
].
Despite
these
advancements,
the
application
of
ML
in
ADMET
prediction
remains
limited
by
the
scope
of
publicly
available
training datasets. Although certain AI models have shown promise
in predicting ADMET and activity properties, a critical challenge lies
in
the
scarcity
of
data
and
the
potential
lack
of
generalizability
of
data-dependent
models.
Furthermore,
these
methods
often
focus
on
the
similarity
of
physicochemical
properties
among
approved
drugs without fully accounting for their behavior within biological
systems,
such
as
permeability
and
clearance
rates.
As
a
result,
a
single
predictive
score
often
fails
to
capture
the
full
complexity
of
drug
properties,
restricting
its
utility
for
guiding
compound
optimization.
4.5.
Accelerating
clinical
trials
Despite
promising
advances
in
systems
biology
and
the
increased
availability
of
high-throughput
biological
data,
the
pharmaceutical
industry
continues
to
face
a
decline
in
R
&
D
ef
fi
-
ciency. Clinical trial failure rates, particularly in oncology and other
disease
areas,
can
reach
as
high
as
95%.
These
high
failure
rates
contribute
signi
fi
cantly
to
the
inef
fi
ciencies
and
costs
of
drug
development: bringing a completely new chemical entity to market
can
take
between
7
and
10
years
of
clinical
trials,
at
a
capitalized
cost ranging from $1.46 billion to $2.56 billion. The
fi
nancial losses
per
failed
clinical
trial
can
range
from
$800
million
to
$1.4
billion,
accounting
not
only
for
the
trial
costs
but
also
for
the
losses
in
preclinical
development.
Applying
AI
technology
to
key
steps
in
clinical
trial
design
offers
the
potential
to
improve
patient
strati
fi
-
cation, enhance recruitment ef
fi
ciency, and ultimately increase the
likelihood
of
trial
success
[
199
e
202
].
In
vivo
studies,
which
account
for
over
75%
of
the
cost
of
developing
new
chemical
entities,
dominate
drug
development
costs.
Therefore,
improvements
in
computational
methods
made
early
in
drug
development,
although
valuable,
have
a
limited
impact
on
reducing
overall
development
expenses.
The
fi
nancial
impact
of
failure
in
phase
III
clinical
trials,
which
involve
large
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
10
patient populations, can be catastrophic. Ideally, AI models used to
predict
late-stage
trial
outcomes,
such
as
in
silico
clinical
trials
(ISCT),
could
signi
fi
cantly
reduce
these
costs
while
improving
overall
success
rates
[
203
].
The
concept
of
Virtual
Physiological
Human
(VPH),
fi
rst
described
in
a
white
paper
in
2005,
envisions
the development of patient-speci
fi
c computer models that support
clinical
decision-making
by
forming
virtual
patient
groups
to
test
the safety and ef
fi
cacy of new drugs and medical devices [
204
,
205
].
These
virtual patient groups
could complement traditional
clinical
trials
by
reducing
the
number
of
patients
required
and
increasing
the
statistical
power
of
the
results,
as
well
as
suggesting
clinical
decisions.
ISCT
typically
integrates
physiological
and
pathological
data
across
different
spatial
and
temporal
scales
to
generate
patient-speci
fi
c predictions. These predictions can inform decisions
regarding
diagnosis,
prognosis,
dose
selection,
and
the
identi
fi
ca-
tion
of
appropriate
patient
groups
[
206
].
However,
using
ISCT
to
reduce or partially replace
in vivo
experiments presents signi
fi
cant
challenges. One key hurdle is the inherent complexity of accurately
and
quantitatively
modeling
organisms.
Without
addressing
these
complexities,
clinical
trials
alone
may
fail
to
provide
suf
fi
cient
structural
and
design
insights
to
explain
the
failure
of
drug
candi-
dates.
Furthermore,
the
reliability
of
ISCT-based
predictions
still
requires
validation
[
207
].
As
a
result,
current
AI
technologies
are
primarily focused on improving clinical trial success by intervening
in
several
critical
areas:
linking
patient
genetic
data,
electronic
health
records
(EHRs),
medical
literature,
and
clinical
trial
data-
bases
to
predict
clinical
toxicity
and
trial
success;
improving
trial
design; assisting with patient-trial matching and recruitment; and
monitoring
patient
adherence
during
trials
[
208
,
209
].
4.5.1.
Prediction
of
clinical
trial
results
DL
models,
when
applied
to
the
analysis
of
drug
response
and
side
effects,
offer
signi
fi
cant
potential
for
predicting
the
outcomes
of
phase
I/II
clinical
trials.
By
improving
the
prediction
of
clinical
trial
success,
these
models
help
optimize
the
drug
development
process
[
210
,
211
].
A
major cause
of
clinical
trial
failures
is
toxicity,
which
can
often
be
predicted
through
computational
models.
For
example, ProCTOR, a model designed to predict toxicity outcomes,
combines
chemical features of drugs with target-based features to
distinguish
between
FDA-approved
drugs
and
drugs
that
failed
in
clinical
trials
due
to
toxicity.
Using
a
48-feature
set,
including
10
molecular
features,
34
target-related
features
(e.g.,
target
tissue
selectivity),
and
4
drug-like
rules,
ProCTOR
constructs
an
RF
clas-
si
fi
er
that
directly
predicts
the
likelihood
of
a
drug
being
toxic
in
clinical
trials
[
212
].
In
addition
to
toxicity,
a
signi
fi
cant
proportion
of
clinical
trials
fail
for
reasons
other
than
safety,
such
as
ef
fi
cacy,
strategic,
and
fi
nancial
factors.
Ef
fi
cacy
prediction
remains
highly
complex, but combining
in vitro
cellular models with data on drug
side effects can help forecast the success or failure of clinical trials.
Insilico Medicine has developed a DNN based on pathway analysis
techniques
to
predict
drug
side
effects
by analyzing
the
transcrip-
tional
changes
drugs
induce
in
cell
lines.
This
network,
built
using
transcriptomic
data
from
drug-induced
perturbations
in
cell
cul-
tures
and
pathway
activation
scores
generated
by
the
iPANDA
al-
gorithm,
predicts
clinical
trial
outcomes
for
46
side
effects.
By
leveraging
these
data-driven
approaches,
it
is
possible
to
predict
the likelihood of success or failure in clinical trials more effectively
[
213
].
4.5.2.
Clinical
trial
design
AI
technologies
are
increasingly
being
applied
to
enhance
clinical
trial
design,
patient
strati
fi
cation,
recruitment,
and
monitoring,
which
ultimately
improves
trial
ef
fi
ciency
and
suc-
cess
rates.
One
example
is
the
collaboration
between
Johns
Hop-
kins University and the National Cancer Institute (NCI) to improve
clinical
trial
design
for
head
and
neck
squamous
cell
carcinoma
(HNSCC). They utilized the iPANDA pathway analysis algorithm to
study transcriptomic data from 359 oral squamous cell carcinoma
(OSCC)
samples
and
86
white
spot
samples
(precancerous
le-
sions).
This
analysis
identi
fi
ed
differentially
dysregulated
path-
ways
between
tumors
and
normal
tissues,
providing
valuable
insights
into
the
complex
signaling
networks
underlying
HNSCC
and
paving
the
way
for
novel
preventive,
diagnostic,
and
thera-
peutic
strategies
[
214
].
The
advent
of
immune
checkpoint
in-
hibitors
in
the
treatment
of
HNSCC
has
created
a
need
for
more
reliable
characterization
of
the
tumor
microenvironment,
partic-
ularly signaling pathways and genetic alterations associated with
CD8
þ
T
cell
in
fi
ltration.
In
their
study,
researchers
used
RNA
sequencing and 10 chemokine signatures to classify patients with
HNSCC
into
subgroups
with
high
and
low
CD8
þ
T cell
in
fi
ltration
(TCIP-H
and
TCIP-L,
respectively).
iPANDA
was
then
applied
to
analyze differences in signaling pathways, somatic mutations, and
copy
number
aberrations
between
these
subgroups.
The
fi
ndings
revealed
that
TCIP-H
tumors
are
rich
in
immune
checkpoint
molecules,
making
them
promising
candidates
for
combination
immunotherapy.
This
work
provides
a
rationale
for
designing
more
effective
immunotherapy strategies
for
HNSCC
[
215
].
BERG,
an
AI
biotechnology
company,
has
developed
Interroga-
tive
Biology,
a
platform
that
utilizes
Bayesian
AI
analysis
to
inte-
grate
multi-omics
molecular
pro
fi
les
with
clinical
health
data,
creating causal inference networks. This technology has been used
to
evaluate
the
phase
I
clinical
trial
of
BPM31510
in
104
patients
with
advanced
recurrent/refractory
solid
tumors
[
27
].
For
the
fi
rst
time,
patient
tissue
samples
and
biological
fl
uids
were
collected
longitudinally,
allowing
for
pan-omics
analysis
at
multiple
time
points.
This
approach
provided
valuable
biological
insights
into
BPM31510's
mechanism
of
action,
con
fi
rming
that
Interrogative
Biology
can
be
used
to
assess
disease
biomarkers
and
develop
actionable drug adverse event management plans. Such plans could
include
excluding
patient
subgroups
that
may
experience
adverse
reactions
or
integrating
preventive
interventions
into
clinical
trial
designs
[
26
].
NLP techniques have also been employed to extract information
from
electronic
medical
records
(EMRs)
to
match
patients
with
suitable
clinical
trials.
IBM
Watson
has
developed
a
clinical
trial
matching
system
that
uses
both
structured
and
unstructured
pa-
tient data from EMRs. This system creates detailed clinical
pro
fi
les
for
patients
and
compares
them
with
the
eligibility
criteria
of
available
clinical
trials,
facilitating
the
optimization
of
clinical
trial
protocols and
improving
patient
recruitment ef
fi
ciency
[
216
,
217
].
Despite
these
advancements,
the
limited
data
available
from
clinical
trials
and
EMRs
does
not
fully
capture
the
complexity
of
biological
systems,
and
the
lack
of
interpretability
raises
concerns
about
the
reliability
of
these
systems
and
poses
ethical
risks.
This
has
led
to
the
growing
importance
of
systems
biology
approaches
that
leverage
medical
data
to
offer
deeper
biological
insights
into
drug
candidates'
mechanisms
of
action.
For
example,
Insilico
Medicine's
iPANDA
algorithm
used
a
Microarray
Analysis
Quality
Control (MAQC) dataset derived from paclitaxel-based neoadjuvant
breast
cancer
therapy
to
identify
biologically
relevant
pathway
features.
These
features
were
successfully
used
to
characterize
patients
with
breast
cancer
based
on
their
sensitivity
to
neo-
adjuvant
therapy.
Similarly,
GNS
Healthcare's
Reverse
Engineering
&
Forward Simulation (REFS) platform integrates various sources of
patient
data,
such
as
EMRs,
medical
claims,
next-generation
sequencing,
and
other
histological
data,
to
build
computational
models.
REFS
applies
ML
to
uncover
hidden
drivers
of
cancer
pro-
gression
and
drug
response,
helping
to
identify
new
biomarkers
and
targets
for
disease
and
enabling
more
accurate
patient
strati-
fi
cation
[
218
e
220
].
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
11
To
improve
patient
adherence
in
clinical
trials,
innovative
technologies
are
being
employed
to
ensure
more
reliable
moni-
toring.
Traditional
methods
of
tracking
adherence,
like
pill
counts
or
self-reported
data,
are
prone
to
manipulation
and
inaccuracies.
AbbVie has implemented AI-powered facial and image recognition
algorithms
through
the
AiCure
mobile
SaaS
platform,
which
re-
quires
patients
to
record
a
video
of
themselves
swallowing
pills.
The
AI
system
then
veri
fi
es
that
the
correct
person
has
taken
the
prescribed
medication.
In
a
study
of
patients
with
schizophrenia,
adherence
increased
from
50%
to
90%
within
six
months,
demon-
strating
the
effectiveness
of
this
AI-driven
monitoring
approach
[
221
e
223
].
4.5.3.
Drug
redirection
The
process
of
bringing
new
chemical
entities
to
market
en-
counters signi
fi
cant hurdles in terms of development time and cost.
However,
discovering
new
indications
for
an
existing
drug
can
substantially
reduce
these
costs
by
repurposing
it
for
different
diseases. Drug repositioning, or drug repurposing, is a strategy that
identi
fi
es
novel
therapeutic
applications for an
approved or
inves-
tigational
drug
outside
its
original
indication
[
224
].
This
approach
allows
repositioned
drugs
to
enter
phases
II
and
III
clinical
trials
more
rapidly,
with
signi
fi
cantly
reduced
development
costs,
as
pharmacokinetic,
pharmacodynamic,
and
toxicity
pro
fi
les
are
already
established
from
earlier
preclinical
and
phase
I
studies.
Historically, successful drug repositioning has often stemmed from
insights
into
drug
pharmacology
or
retrospective
clinical
observa-
tions.
For
instance,
sildena
fi
l
citrate,
initially
developed
as
an
antihypertensive,
was
later
repurposed
by
P
fi
zer
for
erectile
dysfunction
based
on
clinical
fi
ndings,
and
thalidomide's
use
for
erythema
nodosum
leprosum
(ENL)
and
multiple
myeloma
arose
from serendipitous discovery [
225
]. With the advancement of AIDD
methodologies,
such
serendipitous
successes
can
now
be
more
systematically
identi
fi
ed
and
traced.
By
integrating
systems
biology
with
NLP
techniques,
the
study
of off-label drug use has become a preferred retargeting strategy for
pharmaceutical
companies.
This
approach
leverages
large-scale
histological
data
and
EHRs
from
patients
to
identify
new
drug
in-
dications
[
226
].
BenevolentAI's
cognitive
system,
JACS,
utilizes
AI
tools
and
biomedical
knowledge
graphs
to
uncover
novel
connec-
tions within vast, unstructured datasets, such as disease, drug, and
clinical trial information, enabling drug repurposing and facilitating
the
discovery
of
valuable
new
indications.
In
collaboration
with
BenevolentAI,
Johnson
&
Johnson
entered
an
exclusive
licensing
agreement
for
clinical-stage
candidates,
redeveloping
bavisant,
a
histamine
H3
receptor
inverse
agonist,
originally
intended
for
attention
de
fi
cit
hyperactivity
disorder,
for
the
treatment
of
extreme
daytime
sleepiness
in
Parkinson's
disease,
with
phase
II
clinical
trials
underway
[
227
,
228
].
In
February
2020,
following
the
World
Health
Organization's
declaration
of
COVID-19
as
a
global
health emergency, BenevolentAI employed knowledge mapping to
rapidly
identify
baricitinib,
initially
developed
by
Eli
Lilly
for
rheumatoid
arthritis,
as
a
potential
treatment
for
COVID-19.
Simi-
larly,
TwoXAR's DUMA
™
platform mined
multi-omic data, protein
interactions,
chemical
structures,
and
clinical
data
to
explore
new
uses
for
existing
drugs,
identifying
exenatide
and
olopatadine
as
more effective treatments in animal models of rheumatoid arthritis
[
229
e
231
].
In
summary,
AI
technology
has
generated
numerous
impactful
applications in the pharmaceutical industry. In drug development, AI
can analyze patterns in chemical reactions, identify potential targets,
design
and
screen
candidate
molecules,
and
predict
drug
kinetics
and adverse reactions, all of which contribute to shortening the drug
development
cycle.
However,
challenges
remain
in
predicting
adverse
reactions
and
interactions,
such
as
data
quality
issues
and
low
prediction
accuracy.
Additionally,
AI-generated
drugs
have
not
yet reached the market, and thus, the practical effectiveness of AI in
drug development remains unproven.
5.
The
application
challenges
of
AI
in
the
pharmaceutical
industry
The integration of new technologies, such as digitization in drug
discovery and AI, has long been recognized as an irreversible trend.
However,
as
outlined
above,
it
is
essential
to
acknowledge
the
signi
fi
cant variability in access to various resources, which leads to
differing
levels
of
maturity
in
AI-driven
macromolecular
drug
development
across
different
sectors.
Given
these
disparities,
it
is
clear
that,
at
this
stage
of
highly
uneven
data
distribution
and
un-
derdeveloped
data-sharing
models,
AI-focused
macromolecular
drug
discovery
companies
must
prioritize
the
development
of
robust
data
asset
production
capabilities
to
establish
genuine
dif-
ferentiation.
Potential
areas
for
data
production
include
antibody
screening
(e.g.,
single
B-cell
analysis),
target
protein-function
re-
lationships (e.g., proteomics), target epitope structures, and peptide
structures. The data production platform must be both unique and
directly aligned
with
drug
development
needs.
In
terms
of
data
sharing,
federated
learning
presents
a
more
feasible
solution
compared
to
the
complexities
of
establishing
or
participating
in
a
decentralized
autonomous
organization
(DAO).
The
formation
of
a
data
federation
will
likely
center
around
data
standardization
and
the
level
of
digitalization
within
the
partici-
pating
companies.
Collaborative
efforts
between
two
or
three
companies will be easier to implement in practice, with the barriers
at the data level, driven by data volume or creation methods, being
more reliable
and
sustainable.
In
drug
development,
the
fragmentation
of
data
into
isolated
silos has led to substantial inef
fi
ciencies. One promising solution is
federated learning, which enables
“
cooperative prosperity.
”
In this
model,
participants
are
autonomous
entities,
and
organizations
incentivize
them
to
join
through
effective
incentive
and
bene
fi
t-
sharing
mechanisms,
a
feature
absent
in
traditional
ML.
Federated
learning
allows
multiple
participants
to
collaboratively
train
a
model without sharing their raw data, thereby utilizing distributed
data
from
various
sources
while
protecting
sensitive
information,
such
as
proprietary
and
con
fi
dential
data.
This
decentralized
paradigm
is
expected
to
signi
fi
cantly
enhance
the
success
rate
of
AIDD. Federated learning not only facilitates the collective training
of
a
global
model
using
participants'
datasets
but
also
supports
personalized
models
for
each
participant.
Personalized
federated
learning
recognizes
the unique
characteristics
of
each data source,
akin to creating models tailored for different demographic
groups,
such
as
the
elderly
or
children.
By
incorporating
local
data
char-
acteristics,
these
personalized
models
can
improve
prediction
ac-
curacy,
which
is
particularly
bene
fi
cial
in
drug
development
for
making
more
accurate,
participant-speci
fi
c
predictions.
Federated
transfer learning, a technique that further extends the feature space
and sample size, plays a pivotal role in this context. From a broader
perspective,
federated
learning
can
be
categorized
into
horizontal
and
vertical
schemes.
Horizontal
federated
learning
applies
when
participants
share the
same
feature
space,
such
as
molecular
ECFP
fi
ngerprints.
Vertical
federated
learning,
on
the
other
hand,
caters
to
participants
with
distinct
types
of
input
features.
Combining
these
two
approaches,
horizontal
and
vertical
federated
learning,
into
federated
transfer
learning
enhances
the
ability
to
integrate
data
with
shared
and
proprietary
features
from
multiple
parties.
This
combination
allows
for
the
expansion
of
both
feature
and
sample
spaces,
ultimately
improving
the
model's
robustness.
For
instance,
predicting
clinical
outcomes
for
candidate
drugs
often
requires
the
integration
of
diverse
data
from
pharmaceutical
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
12
companies, hospitals, and patients. Federated transfer learning can
enable the pooling of this data in a way that preserves privacy and
proprietary
information,
generating
signi
fi
cant
value
for
all
stake-
holders
involved.
Federated
learning
represents
a
promising
approach
to
drug
discovery,
offering
a
mechanism
to
leverage
con
fi
dential
datasets
through secure, distributed training. This method addresses a critical
challenge in the
fi
eld, where the availability of large, diverse datasets
has
historically
been
limited.
By
facilitating
the
integration
of
data
across
multiple
institutions,
federated
learning
has
the
potential
to
enhance predictive models, which often operate in narrow contexts.
The security and privacy features inherent in federated learning are
particularly
advantageous
for
drug
discovery,
especially
when
handling
sensitive
data
such
as
genetic
information.
This
approach
can incentivize institutions to share their data, which is essential for
creating
robust,
generalizable
models.
As
data
availability
expands,
the
“
small data
”
problem, commonly encountered in drug discovery,
can be mitigated, leading to improved predictive accuracy and more
insightful recommendations.
However, protecting intellectual property rights associated with
algorithms
remains
a
signi
fi
cant
challenge,
with
barriers
to
entry
being
relatively
short-lived.
In
the
context
of
combining
wet
and
dry
experiments,
AI-driven
macromolecular
drug
development
companies
must
possess
strong
algorithm
development
capabil-
ities,
particularly
on
the
application
side.
Innovation
in
this
area
should be viewed as a continuous, systematic process that requires
iterative
updates.
The
output
should
be
user-friendly
for
life
sci-
entists,
ensuring
seamless
integration
with
wet
experiments.
Key
considerations
include
whether
the
predictions
made
by
algo-
rithms
are
meaningful,
interpretable,
and
capable
of
providing
corrective
feedback
based
on
observed
outcomes.
In
terms
of
algorithmic
development
and
computational
support,
collabora-
tion
among
companies
and
a
more
open
approach
to
re
fi
ning
ca-
pabilities
is
essential.
Regarding
business
models,
success
in
the
drug
development
sector
hinges
on
a
company's
deep
understanding
of
the
drug
development
process.
Without
this
expertise,
contract
research
organizations
(CROs)
will
struggle
to
standardize
their
services,
limiting their ability to expand offerings and attract new clients. In
drug
development,
wet
experiments,
which
take
months
to
conduct,
typically
consume
far
more
time
than
dry
experiments,
which
are
completed
within
days.
A
lack
of
understanding
of
the
drug development process can dilute the ef
fi
ciency gains from dry
experiments,
as
the
inef
fi
ciencies
of
wet
experiments
counterbal-
ance
the
bene
fi
ts.
From
the
perspective
of
publicly
traded
com-
panies,
direct
involvement
in
drug
development
has
proven
more
valuable
than
offering
CRO
services.
The
primary
reason
for
this
is
that
downstream
customers
often
lack
the
criteria
to
assess
the
quality
of
AI
algorithms
provided
by
CROs,
resulting
in
low
will-
ingness
to pay,
an
issue
that
is
particularly pronounced
in
China.
Drug
development
companies
require
expertise
in
both
life
sciences
and
algorithms,
with
the
essential
task
of
developing
methodologies
that
effectively
bridge
wet
and
dry
experiments.
This
involves
decisions
on
data
set
selection,
model
training
pro-
cesses
that
yield
predictive
results,
and
how
to
align
those
pre-
dictions with project advancement. Achieving success in AI-driven
macromolecular drug development demands a deep integration of
these
capabilities.
To
accomplish
this,
companies
must
master
technologies
that
extend
beyond
traditional
ML
algorithms
and
biopharmaceuticals,
incorporating
knowledge
from
fi
elds
such
as
synthetic
biology
(e.g.,
the
synthesis
of
arti
fi
cial
proteins)
and
en-
gineering
automation
(e.g.,
digital
adaptation
of
laboratory
automation).
Given the current challenges, there is signi
fi
cant demand within
the
pharmaceutical
industry
for
advanced
technologies
capable
of
accelerating
drug
discovery
and
validation
(
Fig.
4
).
Hundreds
of
collaborations
between
pharmaceutical
companies
and
AI
tech
fi
rms have already occurred worldwide, marking a notable shift
in
the
pharmaceutical
sector
from
skepticism
to
interest
in
AI.
How-
ever,
the
question
remains:
how
far
is
this
shift
from
interest
to
trust? AI's potential to reshape the pharmaceutical landscape could
integrate
the
entire
AI
ecosystem
into
pharmaceutical
industry.
This
raises
the
question
of
whether
computational
and
traditional
pharmaceutical
industry
will
eventually
become
parallel
models,
much
like
online
and
of
fl
ine
shopping,
where
online
shopping
serves
as
a
form
of
VS.
The
future
remains
uncertain.
One
of
the
central challenges is whether the vast array of variables in complex
biological
systems
can
be
accurately
quanti
fi
ed
and
analyzed
to
identify
novel
drug
targets
and
better
assess
the
effects
of
drugs.
many
unknowns
remain
to
be
explored,
one
undeniable
truth
is
that AI has the potential to extract value from data across the entire
drug
discovery
cycle.
Although
data
is
not
synonymous
with
sci-
ence,
virtually
all
scienti
fi
c
breakthroughs
are
identi
fi
ed
and
vali-
dated through data. As the volume of data continues to grow, drug
development
data is
evolving
into big data,
and AI is currently the
most
effective
tool
for
managing
and
extracting
insights
from
this
data.
The
rapid
advancement
of
AI
technology
has
cemented
the
digitalization
and
AI
applications
in
drug
discovery
and
develop-
ment
as
an
irreversible
trend.
However,
due
to
varying
challenges
in data acquisition, the maturity of AI-driven macromolecular drug
development
remains
uneven
across
its
stages.
To
differentiate
themselves,
AI
drug
discovery
companies
must
establish
robust
data asset production capabilities. These platforms must be unique
and closely aligned with the drug development process. In terms of
data sharing, federated learning offers a more feasible solution than
establishing
or
participating
in
a
DAO,
as
it
prioritizes
data
stan-
dardization
and
the
digital
maturity
of
participating
entities.
Federated
learning
ensures
data
privacy
through
distributed
training, thereby enhancing the success rate of AIDD while enabling
the
creation
of
personalized
models
for
participants.
Moreover,
AI
drug
development
companies
must
foster
the
ability
to
innovate
algorithmically
and
maintain
an
open,
collaborative
approach
to
further
strengthen
their
capabilities.
Business
model-wise,
direct
involvement
in
drug
development
is
likely
to
generate
greater
corporate
value
compared
to
providing
CRO
services.
Drug
devel-
opment
companies
must
also
cultivate
expertise
in
both
life
sci-
ences
and
algorithms,
developing
methodologies
that
effectively
integrate wet lab
and
dry lab
experiments. Despite the
challenges,
the
demand
for
advanced
technologies
that
facilitate
drug
discov-
ery
and
validation
remains
immense
within
the
pharmaceutical
industry,
evidenced
by
hundreds
of
global
collaborations.
The
attitude
of
the
traditional
pharmaceutical
industry
toward
AI
has
shifted
from
skepticism
to
interest,
though
a
signi
fi
cant
gap
re-
mains between interest and trust. AI's application in drug discovery
will
likely
integrate
the
entire
AI
ecosystem
into
the
pharmaceu-
tical
sector.
However,
whether
computational
and
traditional
pharmaceutical models will evolve into parallel systems remains to
be seen. Ultimately, whether AI can reshape and transform the drug
discovery
process
by
extracting
value
from
data
across
its
entire
cycle will
de
fi
ne
the
future trajectory of
the
industry.
6.
Overcoming
the
challenges
and
potential
future directions
AI,
as
one
of
the
most
advanced
technologies
today,
has
made
signi
fi
cant
progress across various
industries.
However,
despite its
impressive
performance,
AI
still
faces
numerous
limitations
and
challenges
that
hinder
its
deeper
development
and
broader
appli-
cation. A key issue is the reliance on large volumes of high-quality
data. While unsupervised learning and reinforcement learning may
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
13

reduce
the
demand
for
labeled
data
in
certain
cases,
many
AI
ap-
plications still require extensive datasets for training. For instance,
tasks
such
as
image
recognition
and
NLP
typically
necessitate
millions
or
even
billions
of
data
samples
to
achieve
high
perfor-
mance. Data quality is crucial; errors, biases, or gaps within training
data can adversely impact the outputs of AI models. Although data
cleansing and preprocessing are essential for ensuring data quality,
these
processes
are
often
time-consuming
and
complex.
Further-
more,
the
collection
and
use
of
large
datasets
raise
concerns
regarding
privacy
and
security.
Leaks
or
misuse
of
user
data
may
result in severe violations of privacy rights. Consequently, ensuring
privacy protection while utilizing data for AI applications presents a
signi
fi
cant
challenge.
Another
pressing
issue
is
the
transparency
and
interpretability
of
models.
Many
AI
models,
particularly
DL
models,
are
often
regarded
as
“
black
boxes
”
due
to
the
dif
fi
culty
in
explaining
their
internal
mechanisms.
While
these
models
can
provide
accurate
predictions,
understanding
how
they
make
decisions
remains
challenging, especially in high-stakes
fi
elds such as healthcare and
fi
nance.
To
foster
trust
and
accountability,
AI
models
must
be
interpretable. Stakeholders, including users and regulatory bodies,
need to understand how decisions are made to ensure fairness and
reliability.
Although
researchers
are
developing
techniques
to
enhance
transparency
and
interpretability,
this
issue
remains
unresolved.
Bias
and
fairness
are
also
prominent
concerns
within
DL.
AI
systems may inherit or even amplify biases present in their training
data.
For
example,
datasets
that
include
biases
related
to
gender,
race,
or
other
factors
may
lead
to
discriminatory
outcomes.
Even
when
data
appears
neutral,
algorithms
can
introduce
biases
through
inconsistent
processing
of
data
from
different
groups,
potentially disadvantaging certain populations. Ensuring fairness in
AI systems is a complex task that requires stringent control at every
stage,
from
model
design
and
data
collection
to
model
evaluation.
Researchers
and
developers
are
actively
exploring
strategies
to
reduce and eliminate bias, but achieving completely fair AI remains
a
challenging
goal.
Training
complex
AI
models
demands
substantial
computa-
tional
resources
and
time.
For
example,
training
a
large
DL
model
may
take
weeks
and
require
immense
computational
power.
This
presents
a
signi
fi
cant
challenge
for
resource-constrained
in-
stitutions
and
developers.
Additionally,
the
training
and
operation
of
AI
models
consume
considerable
energy,
leading
to high
power
consumption
and
a
substantial
carbon
footprint.
With
the
continued
proliferation
of
AI
applications,
addressing
the
envi-
ronmental
impact
of
AI
and
reducing
energy
consumption
has
become an increasingly important issue. Researchers are striving to
develop
more
ef
fi
cient
algorithms
and
hardware
to
diminish
the
computational
demands
and
energy
consumption
of
AI
systems.
Moreover,
advancements
in
quantum
computing
and
edge
computing
technologies
hold
promise
for
signi
fi
cantly
enhancing
the
ef
fi
ciency
of
AI systems
in
the
future.
Currently, most AI systems fall under the category of narrow AI,
optimized for speci
fi
c tasks. For instance, an AI designed for image
recognition
excels
in
identifying
images
but
cannot
perform
NLP
tasks.
The
realization
of
arti
fi
cial
general
intelligence
(AGI),
an
AI
capable
of
executing
various
tasks
across
different
domains,
re-
mains
a
distant
goal.
Existing
AI
systems
are
limited
in
their
un-
derstanding
and
reasoning
capabilities.
While
they
can
process
extensive
datasets
and
conduct
complex
computations,
they
still
struggle with common-sense reasoning, contextual judgment, and
complex
problem-solving.
This
restricts
their
effectiveness
in
broader applications. AGI would need to possess the ability to learn
autonomously
and
adapt
to
new
environments;
however,
current
AI
systems
face
dif
fi
culties
in
adapting
to
dynamic
and
changing
environments, often requiring retraining or signi
fi
cant adjustments
Fig.
4.
Application
challenges
of
arti
fi
cial
intelligence
(AI)
in
drug
research
and
development
(R
&
D).
QC:
quality
control;
IP:
intellectual
property;
P
&
L:
pro
fi
t
and
loss;
KPIs:
key
performance
indicators;
IT:
information
technology;
NLP:
natural
language
processing;
GAN:
generative
adversarial
network.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
14
to handle
new tasks.
The rapid expansion of AI applications also brings forth a series
of
ethical,
social,
and
legal
challenges,
particularly
in
medical
AI
research.
These
challenges
include
the
ethical
use
of
informed
consent,
ensuring
safety
and
transparency,
mitigating
algorithmic
bias,
protecting
patient
data
privacy,
and
maintaining
the
security
and
ef
fi
cacy
of
AI
technologies.
Issues
related
to
accountability,
intellectual
property,
and
safeguarding
AI
systems
from
cyber
threats further complicate the situation. Healthcare institutions, as
research
channels
and
ethical
overseers,
must
comprehensively
address
these
issues
and
manage
the
inherent
ethical
concerns
within
their
AI
research and
applications.
The swift advancement of AI technologies has led to increasingly
complex and autonomous algorithms, raising the
“
black box
”
issue
where the decision-making processes of algorithms are dif
fi
cult for
humans
to
understand.
This
problem
not
only
hinders
the
inter-
pretability
of
algorithms
but
may
also
result
in
negative
conse-
quences,
such
as
algorithmic
collusion
or
abuse
of
power.
To
address these issues, explainable AI (XAI) has emerged as a critical
solution,
emphasizing
the
transparency
and
interpretability
of
algorithmic
decisions
[
232
].
XAI
enhances
algorithmic
trans-
parency,
enabling
researchers
to
better
understand,
optimize,
and
trust AI systems. In the context of patent disclosures for algorithms,
the
application
of
XAI
becomes
particularly
important.
XAI
assists
patent
examiners
in
understanding
complex
algorithms,
ensuring
that
patent
disclosures
are
adequate
and
transparent.
Patent
pro-
tection
plays
a
pivotal
role
in
promoting
the
interpretability
and
transparency of AIDD models [
233
]. It requires inventors to disclose
technical
details
of
their
algorithms,
including
underlying
princi-
ples
and
decision-making
processes.
These
requirements
urge
in-
ventors to provide suf
fi
cient information, ensuring that others can
understand and potentially replicate their technologies. In the
fi
eld
of AI, especially in drug design, patent protection places signi
fi
cant
emphasis
on
the
interpretability
of
algorithms
to
ensure
that
ex-
aminers
and
the
public
can
comprehend
how
AI
systems
reach
conclusions. For instance, a team by Collins and co-workders [
234
]
at
MIT
utilized
an
interpretable
DL
model
to
identify
a
novel
anti-
biotic
from
over
12
million
compounds,
effectively
targeting
methicillin-resistant
Staphylococcus
aureus
(MRSA).
This
success
story
underscores
the
importance
of
interpretability
in
AIDD,
as
it
enables
researchers
to
understand
how
the
model
makes
pre-
dictions,
thereby
improving
the
design
of
more
effective
thera-
peutic
agents.
The
drug
developed
using
AI
has
progressed
to
clinical
trials,
underscoring
how
interpretability
builds
trust
be-
tween
researchers
and
clinicians
by
clarifying
drug
mechanisms
and
accelerating
development.
The
mandatory
disclosure
re-
quirements
of
patent
protection
signi
fi
cantly
enhance
the
inter-
pretability and transparency of AIDD models. These success stories
indicate
that
interpretability
is
not
only
a
technical
necessity
but
also
a
crucial
factor
in
gaining
societal
trust
and
facilitating
suc-
cessful
commercialization.
In
addition
to
the
aforementioned
content,
new
methods
continue to emerge in the
fi
eld of XAI, bringing new opportunities
for
drug
development.
For
example,
Chemical-explainable
GNN
(ChemXGNN),
as
a
cutting-edge
XAI
approach,
exhibits
unique
advantages
in
drug
discovery
[
235
].
It
combines
the
powerful
molecular
structure
representation
capabilities
of
GNN
with
interpretability
techniques,
enabling
in-depth
analysis
of
drug-
target
interactions
at
the
molecular
level.
When predicting
the
af-
fi
nity
of
a
speci
fi
c
drug
for
a
target,
ChemXGNN
not only
provides
prediction
results
but
also
demonstrates
the
molecular
structural
features that the model focuses on during decision-making through
its
interpretability
module,
such
as
speci
fi
c
chemical
bonds
and
functional
groups.
This
clear
explanation
of
the
molecular
structure-activity
relationship
aids
medicinal
chemists
in
understanding
the
rationale
behind
model
decisions,
allowing
for
more
targeted
optimization
of
drug
molecular
structures
and
accelerating
the
new drug
development
process.
While
AI
offers
signi
fi
cant
advancements
across
various
fi
elds,
its
widespread
adoption
raises
concerns
related
to
ethics,
social
implications,
and
employment.
AI
has
the
potential
to
automate
certain
professions,
which
may
lead
to
job
losses
and
social
chal-
lenges.
Although
AI
also
creates
new
employment
opportunities,
ensuring a balance between supply and demand in the job market
to avoid
mass
unemployment
remains
a
critical
social
issue.
Addi-
tionally, the application of AI in sensitive areas such as military and
surveillance
raises
ethical
and
moral
dilemmas.
For
instance,
the
use
of
automated
weapons
and
mass
surveillance
systems
may
violate
human
rights
and
privacy,
raising
concerns
about
the
responsible use of AI technologies. As AI technology rapidly evolves,
existing
legal
and
regulatory
frameworks
often
struggle
to
keep
pace.
Governments
and
international
organizations
must
collabo-
rate
to
develop
and
implement
effective
laws
and
regulations
to
govern
the
development
of
AI
and
ensure
the
protection
of
public
interests.
The security and privacy of medical data form the foundation for
advancements
in
AI
within
the
pharmaceutical
sector.
To
meet
the
growing
demand
for
medical
data,
the
development
of
advanced
data
protection
technologies
is
crucial.
Simultaneously,
efforts
are
being
made
to
enhance
the
interpretability
and
reliability
of
AI
models,
thereby
bolstering
public
and
regulatory
con
fi
dence
in
AI-driven
pharmaceuticals.
Improving
data
quality
and
sample
representativeness
is
essential
for
enhancing
the
reliability
and
generalization
capabilities
of
AI
models.
This
requires
strict
enforcement of
standards
and
regulations regarding
data
collection
and processing. Clearly de
fi
ning the responsibilities and rights of all
stakeholders,
protecting
intellectual
property,
and
adhering
to
ethical
standards
in
medical
research
and
clinical
trials
are
para-
mount.
These
measures include
formulating
and
enforcing
relevant
laws and regulations, as well as establishing ethical guidelines. Early
collaboration between companies and regulatory agencies is vital for
ensuring
the
compliance
of
AI
models
and
enhancing
their
credi-
bility. Regulatory bodies should establish frameworks and standards
to
facilitate
this
collaboration.
Furthermore,
veri
fi
cation
plans
must
be
developed
to
validate
the
robustness
of
AI
models
and
ongoing
assessments to ensure their effectiveness and reliability.
The
application
of
AI
in
drug
development
holds
immense
promise
but
faces
several
challenges.
Key
issues
include
data
quality and
accessibility,
as
biases
in
training
datasets
may lead
to
unfair
AI
outcomes,
while
data
silos
hinder
the
sharing
and
utili-
zation
of
critical
information.
Transparency
and
interpretability
of
models remain signi
fi
cant obstacles, as DL models typically operate
as
“
black
boxes
”
,
making
it
dif
fi
cult
to
understand
their
internal
workings,
raising
concerns
about
the
reliability
and
fairness
of
AI-
driven
predictions
and
decisions.
Additionally,
the
generalization
and
robustness
of
AI
models
require
further
improvement
to
mitigate risks of over
fi
tting and adversarial attacks. The substantial
computational
resources
and
energy
consumption
required
for
training
and
deploying
complex
AI
models
also
present
environ-
mental
and
logistical
challenges.
Moreover,
the
integration
of
AI
into healthcare introduces ethical, social, and legal issues, including
privacy
and
security,
algorithmic
bias,
fairness,
and
the
establish-
ment
of
accountability
and
regulatory
frameworks.
Despite
these
challenges,
the
potential
of
AI
in
drug
discovery
remains
vast,
necessitating ongoing research, collaboration, and ethical oversight
to ensure
responsible
usage
and
improve patient outcomes.
In
the
realm
of
ADMET
prediction,
new
platforms
continue
to
emerge
with
ongoing
technological
advancements.
For
instance,
admetSAR3.0
[
236
]
has
been
optimized
and
expanded
based
on
existing
ADMET
prediction
platforms.
Compared
to
previous
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
15
versions,
admetSAR3.0
updates
a
substantial
amount
of
experi-
mental
data and constructs more complex and
accurate predictive
models,
signi
fi
cantly
enhancing
the
accuracy
and
reliability
of
predicting
various
ADMET
properties
of
drugs.
It
employs
more
advanced
ML
algorithms
capable
of
better
handling
complex
mo-
lecular
structure
data
and
uncovering
potential
relationships
be-
tween
chemical
structures
and
ADMET
properties.
Additionally,
admetSAR3.0
offers
a
more
user-friendly
interface
and
richer
functionalities,
allowing
predictions
for
individual
ADMET
prop-
erties
while
supporting
joint
predictions
and
analyses
of
multiple
properties,
providing
drug
developers
with
comprehensive
infor-
mation
to
more
accurately
assess
the
drug-likeness
of
candidate
compounds
in
the
early
stages
of
drug
development,
thereby
enhancing
development ef
fi
ciency and
reducing
costs.
The
integration
of
AI
within
the
pharmaceutical
industry
will
increasingly
prioritize
data
privacy
protection,
utilizing
technolo-
gies
such
as
federated
learning
to
securely
integrate
data
across
institutions,
thereby
enhancing
the
predictive
capabilities
of
drug
discovery models. With the advancement of personalized federated
learning,
AI
will
be
able
to
develop
customized
models
tailored
to
individual
participants,
facilitating
more
personalized
drug
devel-
opment
and
treatment
strategies.
Moreover,
it is noteworthy that
large language models are
grad-
ually emerging in the
fi
eld of drug development. With their powerful
language understanding
and
generation
capabilities, large
language
models
are
subtly
transforming
the
traditional
drug
discovery
pro-
cess [
237
]. For instance, during the drug target discovery phase, large
language models can rapidly analyze and comprehend vast amounts
of biomedical literature, extracting potential drug target information.
Previously,
researchers
needed
to
spend
considerable
time
review-
ing
literature
to
identify
disease-relevant
potential
targets,
while
large
language
models
can process
and
synthesize
this
information
in a short period, uncovering new target associations and improving
target
discovery
ef
fi
ciency
[
238
].
In
the
drug
design
phase,
large
language
models
can
understand
and
generate
textual
descriptions
related
to
chemical structures,
assisting
in the
design of
novel
drug
molecules.
They
can
propose
molecular
structures
with
potential
activity
based
on
known
drug
activity
and
structural
relationships,
providing
medicinal
chemists
with
more
design
ideas
and
broad-
ening the possibilities for drug design, thus enhancing the ef
fi
ciency
and
accuracy
of
drug
design.
Additionally,
large
language
models
can
participate
in
the
design
and
evaluation
of
clinical
trials,
providing
suggestions
for
optimizing
clinical
trial
protocols
by
analyzing
extensive
clinical
data
and
research
literature,
helping
to
identify
more
reasonable
trial
endpoints,
sample
sizes,
and
patient
inclusion criteria, thereby improving the quality and success
rate of
clinical trials.
7.
Conclusion
The pharmaceutical industry is currently experiencing exponen-
tial growth in data, and the most effective AI approaches in modeling
do not rely solely on pure AI processes. In fact, the synergy between
humans
and
AI
often
surpasses
the
capabilities
of
either
indepen-
dently.
Much
like
in
the
realm
of
chess,
where
human-machine
collaboration
can
outperform
either
humans
or
computers
acting
alone,
the
integration
of
human
insights
with
AI
technologies
can
yield
superior
outcomes.
AI
methodologies
require
systematic
or-
ganization
and
development
and
attention,
exploration,
and
exper-
imental
application
across
various
fi
elds
can
accelerate
the
maturation
and
innovation
of
AI
technologies.
As
the
cycle
of
“
big
data
/
more precise models
/
better drugs
/
more and better data
”
gradually
matures
in
practice,
AI-driven
advancements
in
pharma-
ceuticals
will be
signi
fi
cantly accelerated.
However,
the widespread
application
and
integration
of
any
technology
require
time,
and
its
development
proceeds
in
a
wave-like
manner.
Before
AI
and
data-
driven
pharmaceutical
models
can
fully
realize
their
potential,
further
exploration
and
practical
application
are
necessary.
In
sum-
mary,
this
review
aims
to
promote
the
application
of
AI
in
drug
discovery,
address
current
challenges,
encourage
collaborative
ef-
forts
among
stakeholders,
ensure
the
protection
of
intellectual
property,
and
outline
a
future
blueprint
in which
AI
plays
a
crucial
role
in
advancing
pharmaceutical
research.
It
calls
upon
the
phar-
maceutical
community
to
leverage
the
powerful
capabilities
of
AI,
with the ultimate goal of improving patient outcomes through more
ef
fi
cient and effective drug discovery.
It
is
noteworthy
that
despite
the
tremendous
potential
of
AI
in
the
pharmaceutical
sector,
several
limitations
persist.
Data
bias
is-
sues
are
particularly
prominent.
The
training
datasets
may exhibit
various
biases,
such
as
sample
selection
bias
and
measurement
bias.
These
biases
can
result
in
unfair
outcomes
from
AI
models,
affecting
their
predictive
accuracy
and
reliability.
For
instance,
in
drug target prediction, if
the training dataset contains an excess of
target data related to a particular disease while having insuf
fi
cient
data
for
others,
the
model
may
predict
targets
for
the
data-rich
disease
more
accurately,
while
predictions
for
the
data-scarce
dis-
eases may be biased, thereby impacting the speci
fi
city and ef
fi
cacy
of
new drug
development.
Model
interpretability
presents
another
signi
fi
cant
challenge.
Many AI models, especially DL models, are often regarded as
“
black
boxes.
”
For
example,
while
DL
models
used
for
drug
design
can
predict the activity of compounds, they struggle to explain the basis
of
their
predictions
and
the
internal
decision-making
processes.
This
lack
of
transparency
makes
it
dif
fi
cult
for
researchers
to
comprehend model behavior and assess the reliability and safety of
these models, limiting their application and promotion in practical
drug
development.
Ethical issues cannot be overlooked either. The application of AI
in the pharmaceutical
fi
eld involves the handling of vast amounts of
patient
data,
which
includes
personal
privacy
information.
Any
leakage
or
improper
use
of
this
data
could
severely
infringe
on
patients'
privacy
rights.
Additionally,
in
clinical
trials,
when
AI
models
are
employed
for
patient
strati
fi
cation
and
selection,
any
biases
in
the
models
may
lead
to
the
unfair
exclusion
of
certain
patient
groups
from
trials,
thereby
affecting
the
fairness
and
accessibility
of
drug
development.
Furthermore,
disputes
exist
regarding
the
ownership
of
patents
generated
by
AI
and
the
delineation
of
responsibilities,
necessitating
further
legal
and
ethical
guidelines
for clari
fi
cation.
In
conclusion,
while
the
development
prospects
for
AI
in
the
pharmaceutical
fi
eld
are
promising,
achieving
its
widespread
application
and
in-depth
development
necessitates
attention
to
and resolution of these limitations. Moving forward, there is a need
to
enhance
data
quality
management,
develop
interpretable
AI
models,
and
establish
robust
ethical
and
legal
frameworks
to
pro-
mote
the
healthy
and
sustainable
development
of
AI
in
the
phar-
maceutical
sector.
CRediT authorship
contribution
statement
Chen
Fu:
Writing
e
original
draft,
Software,
Project
adminis-
tration,
Funding
acquisition,
Data
curation,
Conceptualization.
Qiuchen
Chen:
Writing
e
review
&
editing,
Visualization,
Valida-
tion,
Supervision,
Resources,
Project
administration,
Funding
acquisition,
Formal analysis,
Conceptualization.
Declaration
of
competing
interest
The
authors
declare
that
there
are
no
con
fl
icts
of
interest.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
16
Acknowledgments
This
study
was
supported
by
the
National
Natural
Science
Foundation
of
China
(Grant
No.:
82304564)
and
the
Liaoning
Province Education Department Scienti
fi
c Research Funding Project
(Grant No.: LJKZ0777).
References
[1]
G.E.
Gignac,
E.T.
Szodorai,
De
fi
ning
intelligence:
Bridging
the
gap
between
human
and
arti
fi
cial
perspectives,
Intelligence
104
(2024),
101832
.
[2]
T.P.
Theodore
Armand,
K.A.
Nfor,
J.I.
Kim,
et
al.,
Applications
of
arti
fi
cial
in-
telligence,
machine
learning,
and
deep
learning
in
nutrition:
A
systematic
review,
Nutrients
16
(2024),
1073
.
[3]
K.
Sutiene,
P.
Schwendner,
C.
Sipos,
et
al.,
Enhancing
portfolio
management
using
arti
fi
cial
intelligence:
Literature
review,
Front.
Artif.
Intell.
7
(2024),
1371502
.
[4]
H.
Ding,
J.
Tian, W.
Yu,
et
al.,
The
application
of
arti
fi
cial
intelligence
and big
data
in
the
food
industry,
Foods
12
(2023),
4511
.
[5]
A.
Sierpe,
R.W.
Yen,
G.
Stevens,
et
al.,
Agenda-setting
in
the
clinical
encounter:
A
systematic
review
protocol,
PLoS
One
19
(2024),
e0312613
.
[6]
D.
Meng,
S.
Zhang,
Y.
Huang,
et
al.,
Application
of
AI
in
biological
age
pre-
diction,
Curr.
Opin.
Struct.
Biol.
85
(2024),
102777
.
[7]
M.T.R.
Hamid,
N.A.
Mumin,
S.A.
Hamid,
et
al.,
Application
of
arti
fi
cial
intel-
ligence (AI) system in opportunistic screening and diagnostic population in a
middle-income
nation,
Curr.
Med.
Imaging
20
(2024),
e15734056280191
.
[8]
J.
Czech,
M.
Willig,
A.
Beyer,
et
al.,
Learning
to
play
the
chess
variant
crazy-
house
above
world
champion
level
with
deep
neural
networks
and
human
data,
Front.
Artif.
Intell.
3
(2020),
24
.
[9]
R.F.
Service,
AI
tools
set
off
an
explosion
of
designer
proteins,
Science
386
(2024)
260
e
261
.
[10]
E.
Callaway,
Chemistry
Nobel
goes
to
developers
of
AlphaFold
AI
that
pre-
dicts
protein
structures,
Nature
634
(2024)
525
e
526
.
[11]
X.
Zou,
W.
He,
Y.
Huang,
et
al.,
AI-driven
diagnostic
assistance
in
medical
inquiry:
Reinforcement
learning
algorithm
development
and
validation,
J.
Med.
Internet
Res.
26
(2024),
e54616
.
[12]
S.
Pilehvari,
Y.
Morgan,
W.
Peng,
An
analytical
review
on
the
use
of
arti
fi
cial
intelligence
and
machine
learning
in
diagnosis,
prediction,
and
risk
factor
analysis
of
multiple
sclerosis,
Mult.
Scler.
Relat.
Disord.
89
(2024),
105761
.
[13]
R.S.
Huang,
A.
Benour,
J.
Kemppainen,
et
al.,
The
future
of
AI
clinicians:
Assessing the modern standard of
chatbots and their approach to diagnostic
uncertainty,
BMC
Med.
Educ.
24
(2024),
1133
.
[14]
L.
Fu,
S.
Shi,
J.
Yi,
et
al.,
ADMETlab
3.0:
An
updated
comprehensive
online
ADMET
prediction
platform
enhanced
with
broader
coverage,
improved
performance,
API
functionality
and
decision
support,
Nucleic
Acids
Res.
52
(2024)
W422
e
W431
.
[15]
K.
Chakravarty,
V.
Antontsev,
Y.
Bundey,
et
al.,
Driving
success
in
personal-
ized
medicine
through
AI-enabled
computational
modeling,
Drug
Discov.
Today
26
(2021)
1459
e
1465
.
[16]
M.P.
Lourenço,
J.
Hosta
s,
C.
Bellinger,
et
al.,
Reinforcement
learning
for
in
silico
determination of adsorbate
Substrate structures, J. Comput. Chem. 45
(2024)
1289
e
1302
.
[17]
M.
Heinzinger,
B.
Rost,
Arti
fi
cial
intelligence
learns
protein
prediction,
Cold
Spring
Harb.
Perspect.
Biol.
16
(2024),
a041458
.
[18]
Healx,
Healx
receives
IND
and
orphan
drug
designation
for
fragile
X
clinical
trial.
https://healx.ai/ind-fragile-x-clinical-trial/
. (Accessed 22 January 2025).
[19]
Healx,
An
AI
fi
rst:
Deep
genomics
platform
reveals
wilson's
disease
drug
candidate.
https://www.biospace.com/article/deep-genomics-ai-program-
reveals-
fi
rst-drug-candidate-for-wilson-s-disease-/
.
(Accessed
22
January
2025).
[20]
A. Zhavoronkov, Y.A. Ivanenkov, A. Aliper, et al., Deep learning enables rapid
identi
fi
cation
of
potent
DDR1
kinase
inhibitors,
Nat.
Biotechnol.
37
(2019)
1038
e
1040
.
[21]
B. Thompson, N. Petri
c Howe, Alphafold 3.0: The AI protein predictor gets an
upgrade,
Nature
(2024),
https://doi.org/10.1038/d41586-024-01385-x
.
[22]
F.
Ren,
A.
Aliper,
J.
Chen,
et
al.,
A
small-molecule
TNIK
inhibitor
targets
fi
brosis
in
preclinical
and
clinical
models,
Nat.
Biotechnol.
43
(2025)
63
e
75
.
[23]
K.A.
Papavassiliou,
A.A.
So
fi
anidi,
V.A.
Gogou,
et
al.,
The
promise
of
arti
fi
cial
intelligence
in
reshaping
anticancer
drug
development,
Cells
13
(2024),
1709
.
[24]
ClinicalTrail.
gov,
Evaluate
REC-4881
in
patients
with
FAP (TUPELO).
https://
clinicaltrials.gov/study/NCT05552755
.
(Accessed
22
July
2025).
[25]
ClinicalTrail.
gov,
Ef
fi
cacy
and
safety
of
REC-2282
in
patients
with
progres-
sive
neuro
fi
bromatosis
type
2
(NF2)
mutated
meningiomas
(POPLAR-NF2).
https://clinicaltrials.gov/study/NCT05130866
.
(Accessed
22
July
2025).
[26]
ClinicalTrail.
gov,
The
symptomatic
cerebral
cavernous
malformation
trial
of
REC-994
(SYCAMORE).
https://clinicaltrials.gov/study/NCT05085561
.
(Accessed
22
July
2025).
[27]
ClinicalTrail.
gov,
Irofulven
in
AR-targeted
and
docetaxel-pretreated
mCRPC
patients
with
drug
response
predictor
(DRP
®
).
https://clinicaltrials.gov/
study/NCT03643107
.
(Accessed
22
July
2025).
[28]
J.N.
Bodor,
J.
Dowell,
J.
Treat,
et
al.,
Phase
II
trial
of
LP-300
in
combination
with
carboplatin
and
pemetrexed
in
never
smoker
patients
with
relapsed
advanced primary adenocarcinoma of the lung after treatment with tyrosine
kinase
inhibitors,
Clin.
Lung
Cancer
(2025),
https://doi.org/10.1016/
j.cllc.2025.05.013
.
[29]
ClinicalTrail.gov,
Study
of
LP-184
in
patients
with
advanced
solid
tumors.
https://www.clinicaltrials.gov/study/NCT05933265
.
(Accessed
22
January
2025).
[30]
ClinicalTrail
gov,
RLY-1971
in
subjects
with
advanced
or
metastatic
solid
tumors.
https://clinicaltrials.gov/study/NCT04252339
.
(Accessed
22
January
2025).
[31]
H.
Sch
€
onherr,
P.
Ayaz,
A.M.
Taylor,
et
al.,
Discovery
of
lirafugratinib
(RLY-
4008), a highly selective irreversible small-molecule inhibitor of FGFR2, Proc.
Natl.
Acad.
Sci.
U
S
A
121
(2024),
e2317756121
.
[32]
A.
Varkaris,
E.
Pazolli,
H.
Gunaydin,
et
al.,
Discovery
and
clinical
proof-of-
concept
of
RLY-2608,
a
fi
rst-in-class
mutant-selective
allosteric
PI3K
a
in-
hibitor
that
decouples
antitumor
activity
from
hyperinsulinemia,
Cancer
Discov
14
(2024)
240
e
257
.
[33]
ClinicalTrail.
gov,
A
study
of
AC682
in
chinese
patients
with
ER
+
/HER2
locally
advanced
or
metastatic
breast
cancer.
https://www.clinicaltrials.gov/
study/NCT05489679
.
(Accessed
22
January
2025).
[34]
ClinicalTrail. gov, A study of AC176 for the treatment of metastatic castration
resistant
prostate
cancer.
https://www.clinicaltrials.gov/study/
NCT05241613
.
(Accessed
22
July
2025).
[35]
ClinicalTrail.
gov,
A
study
of
BPM31510
with
vitamin
K1
in
subjects
with
newly
diagnosed
glioblastoma
(GB).
https://clinicaltrials.gov/study/
NCT04752813
.
(Accessed
22
July
2025).
[36]
M.E.
Lacouture,
H.
Dion,
S.
Ravipaty,
et
al.,
A
phase
I
safety
study
of
topical
calcitriol (BPM31543) for the prevention of chemotherapy-induced alopecia,
Breast
Cancer
Res.
Treat.
186
(2021)
107
e
114
.
[37]
ClinicalTrail.
gov,
A
phase
2a
study
of
LAM-001
for
the
treatment
of
pul-
monary
hypertension.
https://ctv.veeva.com/study/a-phase-2a-study-of-
lam-001-for-the-treatment-of-ph
.
(Accessed
22
January
2025).
[38]
ClinicalTrail. gov, Study
of safety, tolerability, and biological activity of
LAM-
002A
in
C9ORF72-associated
amyotrophic
lateral
sclerosis.
https://classic.
clinicaltrials.gov/ct2/show/NCT05163886
.
(Accessed
22
January
2025).
[39]
ClinicalTrail. gov, A
fi
rst-in-human PoC study with BEN2293 in patients with
mild
to
moderate
atopic
dermatitis.
https://classic.clinicaltrials.gov/ct2/
show/NCT04737304
.
(Accessed
22
January
2025).
[40]
J.D.
Jones,
L.
Rajachandran,
F.
Yocca,
et
al.,
Sublingual
dexmedetomidine
(BXCL501) reduces opioid withdrawal symptoms: Findings from a multi-site,
phase
1b/2,
randomized,
double-blind,
placebo-controlled
trial,
Am.
J.
Drug
Alcohol.
Abuse
49
(2023)
109
e
122
.
[41]
ClinicalTrail.
gov,
A
trial
of
BXCL701
and
pembrolizumab
in
patients
with
mCRPC
either
small
cell
neuroendocrine
prostate
cancer
or
adenocarcinoma
phenotype.
https://clinicaltrials.gov/study/NCT03910660
.
(Accessed
22
January
2025).
[42]
ClinicalTrail. gov, 3-part study to assess safety, tolerability, pharmacokinetics
and
pharmacodynamics
of
EXS21546.
https://classic.clinicaltrials.gov/ct2/
show/NCT04727138
.
(Accessed
22
January
2025).
[43]
ClinicalTrail. gov, A single arm trial evaluating the ef
fi
cacy and safety of EVX-
01
in
combination
with
pembrolizumab
in
adults
with
unresectable
or
metastatic
melanoma.
https://classic.clinicaltrials.gov/ct2/show/
NCT05309421
.
(Accessed
22
January
2025).
[44]
ClinicalTrail.
gov,
Study
of
adjuvant
immunotherapy
with
EVX-02
and
anti-
PD-1.
https://classic.clinicaltrials.gov/ct2/show/NCT04455503
.
(Accessed
22
January
2025).
[45]
ClinicalTrail. gov, Evaluation of the safety, tolerability, pharmacokinetics, and
pharmacodynamics
of
PHI
101
for
the
treatment
of
AML.
https://classic.
clinicaltrials.gov/ct2/show/NCT04842370
.
(Accessed
22
January
2025).
[46]
S.J. Park,
S.J.
Chang,
D.H.
Suh,
et
al.,
A phase
IA
dose-escalation
study
of
PHI-
101,
a
new
checkpoint
kinase
2
inhibitor,
for
platinum-resistant
recurrent
ovarian
cancer,
BMC
Cancer
22
(2022),
28
.
[47]
ClinicalTrail.
gov,
Study
of
SOM0226
in
familial
amyloid
polyneuropathy.
https://classic.clinicaltrials.gov/ct2/show/NCT02191826
.
(Accessed
22
January
2025).
[48]
ClinicalTrail.
gov,
Ef
fi
cacy
and
safety
on
SOM3355
in
Huntington's
disease
chorea.
https://classic.clinicaltrials.gov/ct2/show/NCT05475483
.
(Accessed
22
January
2025).
[49]
Synapse,
SOM0061.
https://synapse.patsnap.com/drug/eab2e3da54714015a
b730d54d371d6df
.
(Accessed
22
January
2025).
[50]
M. Guerrero,
M. Urbano, E.K. Kim, et
al., Design and synthesis
of
a novel
and
selective
kappa
opioid
receptor
(KOR)
antagonist
(BTRX-335140),
J.
Med.
Chem.
62
(2019)
1761
e
1780
.
[51]
ClinicalTrail.
gov,
Safety
and
tolerability
of
FMT
capsules
in
healthy
volun-
teers.
https://classic.clinicaltrials.gov/ct2/show/NCT05352269
.
(Accessed
22
January
2025).
[52]
M.S.
Lee,
A.R.
Parikh,
D.R.
Spigel,
et
al.,
Preliminary
results
from
ERAS-007
plus
encorafenib
and
cetuximab
(EC)
in
patients
(pts)
with
metastatic
BRAF
V600E
mutated
colorectal
cancer
(CRC)
in
HERKULES-3
study:
A
phase
1b/2
study
of
agents
targeting
the
mitogen-activated
protein
kinase
(MAPK)
pathway
in
pts
with
advanced
gastrointestinal
malignancies
(GI
cancers),
J.
Clin.
Oncol.
41
(2023),
3557
.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
17
[53]
M.
McKean,
E.
Rosen,
M.
Barve,
et
al.,
Abstract
CT184:
Preliminary
dose
escalation
results
of
ERAS-601
in
combination
with
cetuximab
in
FLAGSHP-
1:
A
phase
I
study
of
ERAS-601,
a
potent
and
selective
SHP2
inhibitor,
in
patients with previously treated advanced or metastatic solid tumors, Cancer
Res
83
(2023),
CT184
.
[54]
ClinicalTrail.
gov,
A
study
to
evaluate
the
ef
fi
cacy,
safety,
and
tolerability
of
NDI-034858
in
participants
with
active
psoriatic
arthritis.
https://classic.
clinicaltrials.gov/ct2/show/NCT05153148
.
(Accessed
22
January
2025).
[55]
ClinicalTrail.
gov,
A
study
of
NDI
1150-101
in
patients
with
solid
tumors.
https://classic.clinicaltrials.gov/ct2/show/NCT05128487
.
(Accessed
22
January
2025).
[56]
ClinicalTrail.
gov,
Ef
fi
cacy
and
safety
of
oral
BT-11
in
moderate
to
severe
Crohn's
disease.
https://www.clinicaltrials.gov/study/NCT03870334
.
(Accessed
22
January
2025).
[57]
Safety,
tolerability,
and
pharmacokinetics
of
oral
NX-13
in
active
ulcerative
colitis.
https://ctv.veeva.com/study/safety-tolerability-and-pharmacokinetics-
of-oral-nx-13-in-active-ulcerative-colitis
.
(Accessed
22
January
2025).
[58]
ClinicalTrail.
gov,
Study
of
potential
CYP3A4
induction
by
INDV-2000
in
healthy
adults.
https://www.clinicaltrials.gov/study/NCT05694533
.
(Accessed
22
January
2025).
[59]
ClinicalTrail.
gov,
A
fi
rst
in
human
study
to
evaluate
the
safety,
pharmaco-
kinetics, and pharmacodynamics effects of OC514.
https://www.clinicaltrials.
gov/study/NCT05264038
.
(Accessed
22
January
2025).
[60]
ClinicalTrail.
gov,
A
study
to
assess
the
ef
fi
cacy
and
safety
of
PXT3003
in
Charcot-Marie-Tooth
type
1A.
https://www.clinicaltrials.gov/study/
NCT05092841
.
(Accessed
22
January
2025).
[61]
ClinicalTrail.
gov,
Investigation
of
sulindac
(HLX-0201)
and gaboxadol
(HLX-
0206) in male fragile X syndrome patients aged 13-40 (IMPACT-FXS).
https://
www.clinicaltrials.gov/study/NCT04823052
.
(Accessed
22
January
2025).
[62]
ClinicalTrail.
gov,
LY3819253
(LY-CoV555)
for inpatients
with
COVID-19
(An
ACTIV-3/TICO
treatment
trial).
https://www.clinicaltrials.gov/study/
NCT05780268
.
(Accessed
22
January
2025).
[63]
ClinicalTrail.
gov,
A
phase
1,
evaluate
the
safety,
tolerability,
and
pharma-
cokinetics
of
INS018_055
in healthy subjects.
https://www.clinicaltrials.gov/
study/NCT05154240
.
(Accessed
22
January
2025).
[64]
ClinicalTrail.
gov,
Study
of
SGR-1505
in
mature
B-cell
neoplasms.
https://
clinicaltrials.gov/study/NCT05544019
.
(Accessed
27
July
2025).
[65]
N.
Antonissen,
O.
Tryfonos,
I.B.
Houben,
et
al.,
Arti
fi
cial
intelligence
in
radi-
ology:
173
commercially
available
products
and
their
scienti
fi
c
evidence,
Eur.
Radiol.
(2025),
https://doi.org/10.1007/s00330-025-11830-8
.
[66]
A.R.
Javed,
H.U.
Khan,
M.K.B.
Alomari,
et
al.,
Toward
explainable
AI-
empowered
cognitive
health
assessment,
Front.
Public
Health
11
(2023),
1024195
.
[67]
J.
Jumper,
R.
Evans,
A.
Pritzel,
et
al.,
Applying
and
improving
AlphaFold
at
CASP14,
Proteins
89
(2021)
1711
e
1721
.
[68]
K.H.
Nam,
Evaluation
of
AlphaFold3
for
the
fatty
acids
docking
to
human
fatty
acid-binding
proteins,
J.
Mol.
Graph.
Model.
133
(2024),
108872
.
[69]
R.
Araújo,
L.
Ramalhete,
A.
Viegas,
et
al.,
Simplifying
data
analysis
in
biomedical
research:
An
automated,
user-friendly
tool,
Methods
Protoc
7
(2024),
36
.
[70]
S.
Li,
L.
Peng,
L.
Chen,
et
al.,
Discovery
of
highly
bioactive
peptides
through
hierarchical
structural
information
and
molecular
dynamics
simulations,
J.
Chem.
Inf.
Model.
64
(2024)
8164
e
8175
.
[71]
N.
Almusallam,
F.
Ali,
A.
Masmoudi,
et
al.,
An
omics-driven
computational
model
for
angiogenic
protein
prediction:
Advancing
therapeutic
strategies
with
Ens-deep-AGP,
Int.
J.
Biol.
Macromol.
282
(2024),
136475
.
[72]
P. Lin, H. Li, S. Huang, Deep learning in modeling protein complex structures:
From
contact
prediction
to
end-to-end
approaches,
Curr.
Opin.
Struct.
Biol.
85
(2024),
102789
.
[73]
Z.
Fralish,
D.
Reker,
Taking
a
deep
dive
with
active
learning
for
drug
dis-
covery,
Nat.
Comput.
Sci.
4
(2024)
727
e
728
.
[74]
K.
Kawanishi,
M.
Baba,
R.
Kobayashi,
et
al.,
A
novel
deep
learning
approach
for
analyzing
glomerular
basement
membrane
lesions
in
a
mouse
model
of
X-linked
alport
syndrome,
Am.
J.
Pathol.
195
(2025)
143
e
154
.
[75]
G.
Huang,
Y.
Li,
S.
Jameel,
et
al.,
From
explainable
to
interpretable
deep
learning for natural language processing in healthcare: How far from reality?
Comput.
Struct.
Biotechnol.
J.
24
(2024)
362
e
373
.
[76]
C.
Shi,
T.
Gao,
W.
Lyu,
et
al.,
Deep-learning-driven
discovery
of
SN3-1,
a
potent NLRP3 inhibitor with therapeutic potential for in
fl
ammatory diseases,
J.
Med.
Chem.
67
(2024)
17833
e
17854
.
[77]
Y. Guttman, Z. Kerem, Dietary inhibitors of CYP3A4 are revealed using virtual
screening
by
using
a
new
deep-learning
classi
fi
er,
J.
Agric.
Food
Chem.
70
(2022)
2752
e
2761
.
[78]
L.
Peng,
R.
Hu,
F.
Kong,
et
al.,
Reverse
graph
learning
for
graph
neural
network,
IEEE
Trans.
Neural
Netw.
Learn.
Syst.
35
(2024)
4530
e
4541
.
[79]
J. Li,
Q.
Sun, F.
Zhang, et
al.,
Meta-structure-based graph
attention
networks,
Neural
Netw
171
(2024)
362
e
373
.
[80]
F. Bai, S. Li, H. Li, AI enhances drug discovery and development, Natl. Sci. Rev.
11
(2023),
nwad303
.
[81]
L.
Belenguer,
AI
bias:
Exploring
discriminatory
algorithmic
decision-making
models
and
the
application
of
possible
machine-centric
solutions
adapted
from
the
pharmaceutical
industry,
AI
Ethics
2
(2022)
771
e
787
.
[82]
D.R.
Serrano,
F.C.
Luciano,
B.J.
Anaya,
et
al.,
Arti
fi
cial
intelligence
(AI)
appli-
cations
in
drug
discovery
and
drug
delivery:
Revolutionizing
personalized
medicine,
Pharmaceutics
16
(2024),
1328
.
[83]
S. Bharadwaj, K. Deepika, A. Kumar, et al., Exploring the arti
fi
cial intelligence
and
its
impact
in
pharmaceutical
sciences:
Insights
toward
the
horizons
where technology meets tradition, Chem. Biol. Drug Des. 104 (2024), e14639
.
[84]
T.
James,
H.
Hennig,
Knowledge
graphs
and
their
applications
in
drug
dis-
covery,
Methods
Mol.
Biol.
2716
(2024)
203
e
221
.
[85]
M.M.L.
Araújo,
M.L.O.
Andrade,
G.D.
Oliveira,
et
al.,
From
nature
to
drug:
Overview
and
CADD
approach
of
anacardic
acid
to
propose
their
biological
potential,
Curr.
Top.
Med.
Chem.
(2024),
https://doi.org/10.2174/
0115680266319575240905164313
.
[86]
M.G.R.
Priya,
J.
Manisha,
L.P.M.
Lazar,
et
al.,
Computer-aided
drug
discovery
approaches in the identi
fi
cation of anticancer drugs from natural products: A
review,
Curr.
Comput.
Aided
Drug
Des.
21
(2025)
1
e
14
.
[87]
W.
Xu,
Current
status
of
computational
approaches
for
small
molecule
drug
discovery,
J.
Med.
Chem.
67
(2024)
18633
e
18636
.
[88]
P.
Wall,
T.
Ideker,
Representing
mutations
for
predicting
cancer
drug
response,
Bioinformatics
40
(2024)
i160
e
i168
.
[89]
D. Takaya, Computer-aided drug design using the fragment molecular orbital
method: Current status and future applications for SBDD, Chem. Pharm. Bull.
72
(2024)
781
e
786
.
[90]
D.R.
Martin,
A.
Ajmal,
M.
Meyer,
et
al.,
In
silico
identi
fi
cation
of
phytocon-
stituents
from
Capparis
sepiaria
as
interleukin-1
inhibitors
for
rheumatoid
arthritis:
Molecular
docking,
ADMET
pro
fi
ling,
and
molecular
dynamics
simulation,
Silico
Pharmacol
13
(2025),
106
.
[91]
I. Fatima, A. Rehman, Y. Ding, et al., Breakthroughs in AI and multi-omics for
cancer
drug
discovery:
A
review,
Eur.
J.
Med.
Chem.
280
(2024),
116925
.
[92]
R.
Sharma,
E.
Saghapour,
J.Y.
Chen,
An
NLP-based
technique
to
extract
meaningful
features
from
drug
SMILES,
iScience
27
(2024),
109127
.
[93]
Q.
Zhang,
D.
Mao,
Y.
Tu,
et
al.,
A
new
fi
ngerprint
and
graph
hybrid
neural
network
for
predicting
molecular
properties,
J.
Chem.
Inf.
Model.
64
(2024)
5853
e
5866
.
[94]
G.
Pinto,
M.
Gelzo,
G.
Cernera,
et
al.,
Molecular
fi
ngerprint
by
omics-based
approaches
in saliva
from
patients affected
by
SARS-CoV-2
infection,
J. Mass
Spectrom.
59
(2024),
e5082
.
[95]
J.
Liao,
H.
Chen,
L.
Wei,
et
al.,
GSAML-DTA:
An
interpretable
drug-target
binding af
fi
nity prediction model based on graph neural networks with self-
attention
mechanism
and
mutual
information,
Comput.
Biol.
Med.
150
(2022),
106145
.
[96]
Z.
Zhu,
Z.
Yao,
X.
Zheng,
et
al.,
Drug-target
af
fi
nity
prediction
method
based
on multi-scale information interaction and graph optimization, Comput. Biol.
Med.
167
(2023),
107621
.
[97]
Z.
Zhu,
X.
Zheng,
G.
Qi,
et
al.,
Drug
e
target
binding
af
fi
nity
prediction
model
based
on
multi-scale
diffusion
and
interactive
learning,
Expert
Syst.
Appl.
255
(2024),
124647
.
[98]
Z.
Zhu,
Z.
Yao,
G.
Qi,
et
al.,
Associative
learning
mechanism
for
drug-target
interaction
prediction,
CAAI
Trans.
Intell.
Technol.
8
(2023)
1558
e
1577
.
[99]
R.R.
Syahdi,
S.
Jasial,
I.
Maeda,
et
al.,
Bridging
structure-
and
ligand-based
virtual
screening
through
fragmented
interaction
fi
ngerprint,
ACS
Omega
9
(2024)
38957
e
38969
.
[100]
Z.
Zhang,
Y.
Bian,
A.
Xie,
et
al.,
Can
pretrained
models
really
learn
better
molecular
representations
for
AI-aided
drug
discovery?
J.
Chem.
Inf.
Model.
64
(2024)
2921
e
2930
.
[101]
Y.
Wu,
M.
Li,
J.
Shen,
et
al.,
A
consensual
machine-learning-assisted
QSAR
model
for
effective
bioactivity
prediction
of
xanthine
oxidase
inhibitors
us-
ing
molecular
fi
ngerprints,
Mol.
Divers.
28
(2024)
2033
e
2048
.
[102]
F. S
anchez Rodríguez, A.J. Simpkin, G. Chojnowski, et al., Using deep-learning
predictions
reveals
a
large
number
of
register
errors
in
PDB
depositions,
IUCrJ
11
(2024)
938
e
950
.
[103]
D.
Torodii,
J.B.
Holmes,
P.
Moutzouri,
et
al.,
Crystal
structure
validation
of
verinurad
via
proton-detected
ultra-fast
MAS
NMR
and
machine
learning,
Faraday
Discuss
255
(2025)
143
e
158
.
[104]
P. Mohanraj, V. Raman, S. Ramanathan, Deep learning for Parkinson
’
s disease
diagnosis: A graph neural network (GNN) based classi
fi
cation approach with
graph
wavelet
transform
(GWT)
using
protein-peptide
datasets,
Diagnostics
(Basel)
14
(2024),
2181
.
[105]
P. Duan, C. Zhou, Y. Liu, Dynamic graph representation learning via coupling-
process
model,
IEEE
Trans.
Neural
Netw.
Learn.
Syst.
35
(2024)
12383
e
12395
.
[106]
Z. Fan, L. Chen, X. Wu, et al., Enhancing predictions of drug solubility through
multidimensional
structural
characterization
exploitation,
IEEE
J.
Biomed.
Health
Inform.
29
(2025)
1828
e
1837
.
[107]
X. Gu, J. Liu, Y. Yu, et al., MFD-GDrug: Multimodal feature fusion-based deep
learning
for
GPCR-drug
interaction
prediction,
Methods
223
(2024)
75
e
82
.
[108]
B. Chen, Z. Pan, M. Mou, et al., Is fragment-based graph a better graph-based
molecular
representation
for
drug
design?
A
comparison
study
of
graph-
based
models,
Comput.
Biol.
Med.
169
(2024),
107811
.
[109]
Y.
Yuan,
Z.
Chen,
T.
Feng,
et
al.,
Tripartite
interaction
representation
algo-
rithm
for
crystal
graph
neural
networks,
Sci.
Rep.
14
(2024),
24881
.
[110]
F.
Wang,
C.A.
Barrero,
Multi-omics
analysis
identi
fi
ed
drug
repurposing
targets
for
chronic
obstructive
pulmonary
disease,
Int.
J.
Mol.
Sci.
25
(2024),
11106
.
[111]
S.J.
Nicholls,
A.J.
Nelson,
CETP
inhibitors:
Should
we
continue
to
pursue
this
pathway?
Curr.
Atheroscler.
Rep.
24
(2022)
915
e
923
.
[112]
M.
Fantacuzzi,
R.
Paciotti,
M.
Agamennone,
A
comprehensive
computational
insight into the PD-L1 binding to PD-1 and small molecules, Pharmaceuticals
(Basel)
17
(2024),
316
.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
18
[113]
D.
Gohel,
P.
Zhang,
A.K.
Gupta,
et
al.,
Sildena
fi
l
as
a
candidate
drug
for
Alz-
heimer
’
s
disease:
Real-world
patient
data
observation
and
mechanistic
ob-
servations
from
patient-induced
pluripotent
stem
cell-derived
neurons,
J.
Alzheimers.
Dis.
98
(2024)
643
e
657
.
[114]
Y.
Shi,
L.
Bao,
Y.
Li,
et
al.,
Multi-omics
combined
to
investigate
potential
druggable
therapeutic
targets
for
stroke:
A
systematic
Mendelian
randomi-
zation
study
and
transcriptome
veri
fi
cation,
J.
Affect.
Disord.
366
(2024)
196
e
209
.
[115]
A.J. Clark,
J.W. Lillard
Jr.,
A comprehensive review
of
bioinformatics
tools
for
genomic
biomarker
discovery
driving
precision
oncology,
Genes
(Basel)
15
(2024),
1036
.
[116]
A.
Sau,
L.
Pastika,
E.
Sieliwonczyk,
et
al.,
Arti
fi
cial
intelligence-enabled
elec-
trocardiogram
for
mortality
and
cardiovascular
risk
estimation:
A
model
development and validation study, Lancet Digit. Health 6 (2024) e791
e
e802
.
[117]
G.A. Logsdon, P. Ebert, P.A. Audano, et al., Complex genetic variation in nearly
complete
human
genomes,
Nature
(2025),
https://doi.org/10.1038/s41586-
025-09140-6
.
[118]
I. Jurisica, Explainable biology for improved therapies in precision medicine:
AI
is
not
enough,
Best
Pract.
Res.
Clin.
Rheumatol.
38
(2024),
102006
.
[119]
BioSpace,
BERG
announces
presentation
of
mutliple
interrogative
biology
platform
®
and
outputs
from
multiple
collaborations
at
AACR.
https://www.
biospace.com/article/releases/berg-announces-presentation-of-mutliple-
interrogative-biology-platform-and-outputs-from-multiple-collaborations-
at-aacr/
.
(Accessed
22
January
2025).
[120]
D.
Watanabe,
M.
Hiroshima,
M.
Yasui,
et
al.,
Single
molecule
tracking
based
drug
screening,
Nat.
Commun.
15
(2024),
8975
.
[121]
C. Segú-Verg
es, L. Artigas, M. Coma, et al., Arti
fi
cial intelligence assessment of
the
potential
of
tocilizumab
along
with
corticosteroids
therapy
for
the
management of COVID-19 evoked acute respiratory distress syndrome, PLoS
One
18
(2023),
e0280677
.
[122]
G.
Aghakhanyan,
A.
Galgani,
A.
Vergallo,
et
al.,
Brain
metabolic
correlates
of
Locus
Coeruleus
degeneration
in
Alzheimer
’
s
disease:
A
multimodal
neuro-
imaging
study,
Neurobiol.
Aging
122
(2023)
12
e
21
.
[123]
G.
Yu,
Q.
Ye,
T.
Ruan,
Enhancing
error
detection
on
medical
knowledge
graphs
via
intrinsic
label,
Bioengineering
(Basel)
11
(2024),
225
.
[124]
L. Wang, R. Ding, Y. Zhai, et al., Giant Panda identi
fi
cation, IEEE Trans. Image
Process.
30
(2021)
2837
e
2849
.
[125]
K.
Sujaritwanid,
B.
Suzuki,
E.Y.
Suzuki,
Comparison
of
one
versus
two
maxillary
molars
distalization
with
iPanda:
A
fi
nite
element
analysis,
Prog.
Orthod.
22
(2021),
12
.
[126]
I.V.
Ozerov,
K.V.
Lezhnina,
E.
Izumchenko,
et
al.,
In
silico
Pathway
Activation
Network
Decomposition
Analysis
(iPANDA)
as
a
method
for
biomarker
development,
Nat.
Commun.
7
(2016),
13427
.
[127]
E.
Amiri
Souri,
R.
Laddach,
S.N.
Karagiannis,
et
al.,
Novel
drug-target
in-
teractions
via
link
prediction
and
network
embedding,
BMC
Bioinformatics
23
(2022),
121
.
[128]
F.
Olotu,
E.
Medina-Carmona,
A.
Serrano-Sanchez,
et
al.,
Structure-based
discovery
and
in
vitro
validation
of
inhibitors
of
chloride
intracellular
channel
4
protein,
Comput.
Struct.
Biotechnol.
J.
21
(2023)
688
e
701
.
[129]
A.
Gaurav,
N.
Agrawal,
M.
Al-Nema,
et
al.,
Computational
approaches
in
the
discovery
and
development
of
therapeutic
and
prophylactic
agents
for
viral
diseases,
Curr.
Top.
Med.
Chem.
22
(2022)
2190
e
2206
.
[130]
X.
Wang,
Y.
Shen,
S.
Wang,
et
al.,
PharmMapper
2017
update:
A
web
server
for
potential
drug
target
identi
fi
cation
with
a
comprehensive
target
phar-
macophore
database,
Nucleic
Acids
Res.
45
(2017)
W356
e
W360
.
[131]
M. Sobhy, A. Eletriby, H. Ragy, et al., ACE inhibitors and angiotensin receptor
blockers
for
the
primary
and
secondary
prevention
of
cardiovascular
out-
comes:
Recommendations
from
the
2024
Egyptian
cardiology
expert
consensus
in
collaboration
with
the
CVREP
foundation,
Cardiol.
Ther.
13
(2024)
707
e
736
.
[132]
M.
Li,
R.A.
Soo,
Can
therapeutic
drug
monitoring
of
lorlatinib
help
us
design
the
right
CROWN?
Lung
Cancer
196
(2024),
107965
.
[133]
K.R. Acharya, K.S. Gregory, E.D. Sturrock,
Advances in the structural basis for
angiotensin-1
converting
enzyme
(ACE)
inhibitors,
Biosci.
Rep.
44
(2024),
BSR20240130
.
[134]
J.
Wang,
D.
Beyer,
C.
Vaccarin,
et
al.,
Development
of
radio
fl
uorinated
MLN-
4760
derivatives
for
PET
imaging
of
the
SARS-CoV-2
entry
receptor
ACE2,
Eur.
J.
Nucl.
Med.
Mol.
Imaging
52
(2024)
9
e
21
.
[135]
D.A.
Belinskaia,
N.N.
Shestakova,
K.V.
Samodurova,
et
al.,
Computational
study
of
molecular
mechanism
for
the
involvement
of
human
serum
albu-
min
in
the
renin-angiotensin-aldosterone
system,
Int.
J.
Mol.
Sci.
25
(2024),
10260
.
[136]
L.
Chen,
Q.
Li,
K.F.A.
Nasif,
et
al.,
AI-driven
deep
learning
techniques
in
protein
structure
prediction,
Int.
J.
Mol.
Sci.
25
(2024),
8426
.
[137]
R.
Roy,
H.M.
Al-Hashimi,
AlphaFold3
takes
a
step
toward
decoding
molecular
behaviorand biologicalcomputation, Nat.Struct. Mol.Biol.31(2024)997
e
1000
.
[138]
AlphaFold3
why did
Nature
publish it without its code? Nature 629 (2024),
728
.
[139]
D.
Wu,
R.
Yin,
G.
Chen,
et
al.,
Structural
characterization
and
AlphaFold
modeling
of
human
T
cell
receptor
recognition
of
NRAS
cancer
neoantigens,
Sci.
Adv.
10
(2024),
eadq6150
.
[140]
J.
Wee,
G.-W.
Wei,
Benchmarking
AlphaFold3
’
s
protein-protein
complex
accuracy
and
machine
learning
prediction
reliability
for
binding
free
energy
changes
upon
mutation,
arXiv
(2024),
https://doi.org/10.48550/
arXiv.2406.03979
.
[141]
R.T.
McDonnell,
A.N.
Henderson,
A.H.
Elcock,
Structure
prediction
of
large
RNAs with
AlphaFold3 highlights
its
capabilities and
limitations,
J. Mol. Biol.
436
(2024),
168816
.
[142]
K.M.
Saravanan,
J.F.
Wan,
L.
Dai,
et
al.,
A
deep
learning
based
multi-model
approach
for
predicting
drug-like
chemical
compound
’
s
toxicity,
Methods
226
(2024)
164
e
175
.
[143]
D.
Ang,
C.
Rakovski,
H.S.
Atamian,
De
novo
drug
design
using
transformer-
based machine translation and reinforcement learning of an adaptive Monte
Carlo
tree
search,
Pharmaceuticals
(Basel)
17
(2024),
161
.
[144]
C.
Hasselgren,
T.I.
Oprea,
Arti
fi
cial
intelligence
for
drug
discovery:
Are
we
there
yet,
Annu.
Rev.
Pharmacol.
Toxicol.
64
(2024)
527
e
550
.
[145]
L.
Isigkeit,
T.
H
€
ormann,
E.
Schallmayer,
et
al.,
Automated
design
of
multi-
target
ligands
by
generative
deep
learning,
Nat.
Commun.
15
(2024),
7946
.
[146]
B.A.
Wright,
R.
Sarpong,
Molecular
complexity
as
a
driving
force
for
the
advancement
of
organic
synthesis,
Nat.
Rev.
Chem.
8
(2024)
776
e
792
.
[147]
O.
Wiest,
C.
Bauer,
P.
Helquist,
et
al.,
Finding
relevant
retrosynthetic
dis-
connections
for
stereocontrolled
reactions,
J.
Chem.
Inf.
Model.
64
(2024)
5796
e
5805
.
[148]
E.C. Oelsner, Y. Sun, P.P. Balte, et al., Epidemiologic features of recovery from
SARS-CoV-2
infection,
JAMA
Netw.
Open
7
(2024),
e2417440
.
[149]
S.C.
Fragkouli,
D.
Solanki,
L.J.
Castro,
et
al.,
Synthetic
data:
How
could
it
be
used in infectious disease research? Future Microbiol. 19 (2024) 1439
e
1444
.
[150]
J. Chen, W. Li, P. Hou, et al., An appearance-semantic descriptor with coarse-
to-
fi
ne
matching
for
robust
VPR,
Sensors
(Basel)
24
(2024),
2203
.
[151]
E.
L
opez-Ch
avez,
A.
Garcia-Quiroz,
J.A.I.
Díaz-G
ongora,
et
al.,
Effect
of
gra-
phene
on
the
key
electrical,
optical,
and
magnetic
properties
of
poly-
methylmethacrylate:
A
study
based
on
molecular
modeling,
J.
Mol.
Model.
30
(2024),
375
.
[152]
M.
Wang,
Y.
Liu,
J.
Yuan,
et
al.,
Inter-class
and
inter-domain
semantic
augmentation
for
domain
generalization,
IEEE
Trans.
Image
Process.
33
(2024)
1338
e
1347
.
[153]
M.H.S. Segler, M. Preuss, M.P. Waller, Planning chemical syntheses with deep
neural
networks
and
symbolic
AI,
Nature
555
(2018)
604
e
610
.
[154]
L.F. Salas-Nu
~
nez, A. Barrera-Ocampo, P.A. Caicedo, et al., Machine learning to
predict
enzyme-substrate
interactions
in
elucidation
of
synthesis
pathways:
A
review,
Metabolites
14
(2024),
154
.
[155]
S.R.
Krishnan,
N.
Bung,
R.
Srinivasan,
et
al.,
Target-speci
fi
c
novel
molecules
with their recipe: Incorporating synthesizability in the design process, J. Mol.
Graph.
Model.
129
(2024),
108734
.
[156]
L.
Schoenmaker,
O.J.M.
B
equignon,
W.
Jespers,
et
al.,
UnCorrupt
SMILES:
A
novel
approach
to
de
novo
design,
J.
Cheminf.
15
(2023),
22
.
[157]
U.V.
Ucak,
I.
Ashyrmamatov,
J.
Lee,
Reconstruction
of
lossless
molecular
representations
from
fi
ngerprints,
J.
Cheminform.
15
(2023),
26
.
[158]
M.
Kutsal,
F.
Ucar,
N.
Kati,
Computational
drug
discovery
on
human
immu-
node
fi
ciency
virus
with
a
customized
long
short-term
memory
variational
autoencoder
deep-learning
architecture,
CPT
Pharmacometrics
Syst.
Phar-
macol.
13
(2024)
308
e
316
.
[159]
A.Y. Bande, S. Baday, Accelerating molecular docking using machine learning
methods,
Mol.
Inform.
43
(2024),
e202300167
.
[160]
F.
Li,
J.
Li,
J.
Song,
et
al.,
Identi
fi
cation
of
a
novel
orthonairovirus
from
ticks
and
serological
survey
in
animals
near
China-North
Korea
border,
J.
Med.
Virol.
96
(2024),
e29567
.
[161]
C.
Herbst,
S.
Endres,
R.
Würz,
et
al.,
Assessment
of
fragment
docking
and
scoring
with
the
endothiapepsin
model
system,
Arch.
Pharm.
357
(2024),
2400061
.
[162]
M.
Nehal,
J.
Khatoon,
S.
Akhtar,
et
al.,
Computational
insights
into
inhibiting
EphA2:
Integrating
structure-based
virtual
screening,
docking,
and
molec-
ular
dynamics
simulations
for
small
molecule
discovery,
Cell.
Mol.
Biol.
70
(2024)
16
e
31
.
[163]
J.
Meng,
L.
Zhang,
Z.
He,
et
al.,
Development
of
a
machine
learning-based
target-speci
fi
c
scoring
function
for
structure-based
binding
af
fi
nity
predic-
tion
for
human
dihydroorotate
dehydrogenase
inhibitors,
J.
Comput.
Chem.
46
(2025),
e27510
.
[164]
I.
Mohammadzadeh,
B.
Hajikarimloo,
B.
Niroomand,
et
al.,
Prediction
of
recurrence
after
surgery
for
pituitary
adenoma
using
machine
learning-
based
models:
Systematic
review
and
meta-analysis,
BMC
Endocr
Disord
25
(2025),
158
.
[165]
C.
Shen,
X.
Hu,
J.
Gao,
et
al.,
The
impact
of
cross-docked
poses
on
perfor-
mance
of
machine
learning
classi
fi
er
for
protein-ligand
binding
pose
pre-
diction,
J.
Cheminform.
13
(2021),
81
.
[166]
Y.
Zhang,
M.
Wang,
E.
Zhang,
et
al.,
Arti
fi
cial
intelligence
in
the
screening,
diagnosis,
and
management
of
aortic
stenosis,
Rev.
Cardiovasc.
Med.
25
(2024),
31
.
[167]
X.
Zhang,
C.
Shen,
H.
Zhang,
et
al.,
Advancing
ligand
docking
through
deep
learning:
Challenges
and
prospects
in
virtual
screening,
Acc.
Chem.
Res.
57
(2024)
1500
e
1509
.
[168]
M.
Junaid,
B.
Wang,
W.
Li,
Data-augmented
machine
learning
scoring
func-
tions
for
virtual
screening
of
YTHDF1
m6A
reader
protein,
Comput.
Biol.
Med.
183
(2024),
109268
.
[169]
L.
Gantenbein,
S.E.
Cerminara,
J.T.
Maul,
et
al.,
Arti
fi
cial
intelligence-driven
skin
aging
simulation
as
a
novel
skin
cancer
prevention,
Dermatology
241
(2025)
59
e
71
.
[170]
B.
Clarke,
E.
Holtkamp,
H.
€
Oztürk,
et
al.,
Integration
of
variant
annotations
using
deep
set
networks
boosts
rare
variant
association
testing,
Nat.
Genet.
56
(2024)
2271
e
2280
.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
19
[171]
T.
Hermans,
L.
Smets,
K.
Lemmens,
et
al.,
A
multi-task
and
multi-channel
convolutional
neural
network
for
semi-supervised
neonatal
artefact
detec-
tion,
J.
Neural
Eng.
20
(2023),
026013
.
[172]
Y. Wang, Y. Su, K. Zhao, et al., A deep learning drug screening framework for
integrating
local-global
characteristics:
A
novel
attempt
for
limited
data,
Heliyon
10
(2024),
e34244
.
[173]
P.
Mittal,
D.
Ghanghas,
D.
Sharma,
et
al.,
Thiazolidine-4-one
analogues:
Synthesis,
in-silico
molecular
modeling,
and
in-vivo
estimation
for
anticon-
vulsant
potential,
Cent.
Nerv.
Syst.
Agents
Med.
Chem.
(2024),
https://
doi.org/10.2174/0118715249322920241004113343
.
[174]
N.F.
de
Sousa,
G.D.
Duarte,
C.B.
Moraes,
et
al.,
In
silico
and
in
vitro
studies
of
terpenes
from
the
Fabaceae
family
using
the
phenotypic
screening
model
against
the
SARS-CoV-2
virus,
Pharmaceutics
16
(2024),
912
.
[175]
P.V. Pogodin, E.G. Salina, V.V. Semenov, et al., Ligand-based virtual screening
and
biological
evaluation
of
inhibitors
of
Mycobacterium
tuberculosis
H37Rv,
SAR
QSAR
Environ.
Res.
35
(2024)
53
e
69
.
[176]
K. Mkhayar, O. Daoui, R. Haloui, et al., Ligand-based design of novel quinoline
derivatives
as
potential
anticancer
agents:
An
in-silico
virtual
screening
approach,
Molecules
29
(2024),
426
.
[177]
V.
Kumar,
K.
Jangid,
N.
Kumar,
et
al.,
3D-QSAR-based
pharmacophore
modelling
of
quinazoline
derivatives
for
the
identi
fi
cation
of
acetylcholin-
esterase
inhibitors
through
virtual
screening,
molecular
docking,
molecular
dynamics
and
DFT
studies,
J.
Biomol.
Struct.
Dyn.
43
(2025)
2631
e
2645
.
[178]
K.A.
Krishnan,
S.G.
Valavi,
A.
Joy,
Identi
fi
cation
of
novel
EGFR
inhibitors
for
the
targeted
therapy
of
colorectal
cancer
using
pharmacophore
modelling,
docking,
molecular
dynamic
simulation
and
biological
activity
prediction,
Anticancer
Agents
Med.
Chem.
24
(2024)
263
e
279
.
[179]
A. Karampuri, S. Perugu, A breast cancer-speci
fi
c combinational QSAR model
development
using
machine
learning
and
deep
learning
approaches,
Front.
Bioinform.
3
(2024),
1328262
.
[180]
F. Ghasemi, A. Fassihi, H. P
erez-S
anchez, et al., The role of different sampling
methods
in
improving
biological
activity
prediction
using
deep
belief
network,
J.
Comput.
Chem.
38
(2017)
195
e
203
.
[181]
A.J.
Minnich,
K.
McLoughlin,
M.
Tse,
et
al.,
AMPL:
A
data-driven
modeling
pipeline
for
drug
discovery,
J.
Chem.
Inf.
Model.
60
(2020)
1955
e
1968
.
[182]
J.M. Stokes, K. Yang, K. Swanson, et al., A deep learning approach to antibiotic
discovery,
Cell
180
(2020)
688
e
702.e13
.
[183]
J. Olui
c, K. Nikolic, J. Vucicevic, et al., 3D-QSAR, virtual screening, docking and
design
of
dual
PI3K/mTOR
inhibitors
with
enhanced
antiproliferative
activ-
ity,
Comb.
Chem.
High
Throughput
Screen.
20
(2017)
292
e
303
.
[184]
K. Swanson, P. Walther, J. Leitz, et al., ADMET-AI: A machine learning ADMET
platform
for
evaluation
of
large-scale
chemical
libraries,
Bioinformatics
40
(2024),
btae416
.
[185]
T.T.V.
Tran,
H.
Tayara,
K.T.
Chong,
Recent
studies
of
arti
fi
cial
intelligence
on
in
silico
drug
distribution
prediction,
Int.
J.
Mol.
Sci.
24
(2023),
1815
.
[186]
J.
Jones,
R.D.
Clark,
M.S.
Lawless,
et
al.,
The
AI-driven
drug
design
(AIDD)
platform:
An
interactive
multi-parameter
optimization
system
integrating
molecular
evolution
with
physiologically
based
pharmacokinetic
simula-
tions,
J.
Comput.
Aided
Mol.
Des.
38
(2024),
14
.
[187]
M.B.
Afridi,
H.
Sardar,
G.
Serdaro
glu,
et
al.,
SwissADME
studies
and
Density
Functional
Theory
(DFT)
approaches
of
methyl
substituted
curcumin
de-
rivatives,
Comput.
Biol.
Chem.
112
(2024),
108153
.
[188]
S.
Zhai,
R.
Wang,
J.
Wang,
et
al.,
Curcumol:
A
review
of
its
pharmacology,
pharmacokinetics,
drug
delivery
systems,
structure-activity
relationships,
and
potential
applications,
In
fl
ammopharmacology
32
(2024)
1659
e
1704
.
[189]
I.
Semenyuta,
O.
Golovchenko,
O.
Bahrieieva,
et
al.,
Synthesis,
characteriza-
tion,
in
vitro
anticancer
evaluation,
ADMET
properties,
and
molecular
docking
of
novel
5-sulfanyl
substituted
(thiazol-4-yl)-phosphonium
salts,
ChemMedChem
19
(2024),
e202400205
.
[190]
R.D.
Hernandez,
F.A.F.
Genio,
J.R.
Casanova,
et
al.,
Antiproliferative
activities
and SwissADME predictions of physicochemical properties of carbonyl group-
modi
fi
ed
rotenone
analogues,
ChemistryOpen
13
(2024),
e202300087
.
[191]
S.
Postel-Vinay,
V.K.
Lam,
W.
Ros,
et
al.,
First-in-human
phase
I
study
of
the
OX40 agonist GSK3174998 with or without pembrolizumab in patients with
selected
advanced
solid
tumors
(ENGAGE-1),
J.
Immunother.
Cancer
11
(2023),
e005301
.
[192]
S.
Piperno-Neumann,
M.S.
Carlino,
V.
Boni,
et
al.,
A
phase
I
trial
of
LXS196,
a
protein kinase C (PKC) inhibitor, for metastatic uveal melanoma, Br. J. Cancer
128
(2023)
1040
e
1051
.
[193]
L.
Ortega-Paz,
S.
Giordano,
D.
Capodanno,
et
al.,
Clinical
pharmacokinetics
and pharmacodynamics of CSL112, Clin. Pharmacokinet. 62 (2023) 541
e
558
.
[194]
S.
Bianzano,
A.
Henrich,
L.
Herich,
et
al.,
Ef
fi
cacy
and
safety
of
the
ghrelin-
O
-
acyltransferase inhibitorBI 1356225 in overweight/obesity: Data from twophase
I, randomised, placebo-controlled studies, Metabolism 143 (2023), 155550
.
[195]
E.N.
Feinberg,
E.
Joshi,
V.S.
Pande,
et
al.,
Improvement
in
ADMET
prediction
with
multitask
deep
featurization,
J.
Med.
Chem.
63
(2020)
8835
e
8848
.
[196]
H.
Cai,
H.
Zhang,
D.
Zhao,
et
al.,
FP-GNN:
A
versatile
deep
learning
archi-
tecture
for
enhanced
molecular
property
prediction,
Brief.
Bioinform.
23
(2022),
bbac408
.
[197]
V. Thumma, V. Mallikanti, R. Matta, et al., Design, synthesis, and cytotoxicity
of
ibuprofen-appended
benzoxazole
analogues
against
human
breast
adenocarcinoma,
RSC
Med.
Chem.
15
(2024)
1283
e
1294
.
[198]
B.
Tang,
S.T.
Kramer,
M.
Fang,
et
al.,
A
self-attention
based
message
passing
neural network for predicting molecular lipophilicity and aqueous solubility,
J.
Cheminform.
12
(2020),
15
.
[199]
J.
Veeraraghavan,
C.
De
Angelis,
J.S.
Reis-Filho,
et
al.,
De-escalation
of
treat-
ment
in
HER2-positive
breast
cancer:
Determinants
of
response
and
mech-
anisms
of
resistance,
Breast
34
(2017)
S19
e
S26
.
[200]
H. Lin, R. Li, Z. Liu, et al., Diagnostic ef
fi
cacy and therapeutic decision-making
capacity
of
an
arti
fi
cial
intelligence
platform
for
childhood
cataracts
in
eye
clinics:
A
multicentre
randomized
controlled
trial,
EClinicalMedicine
9
(2019)
52
e
59
.
[201]
N. Li, L. Cui, H. Ma, et al.,
Osteopontin is highly secreted in the cerebrospinal
fl
uid
of
patient
with
posterior
pituitary
involvement
in
Langerhans
cell
Histiocytosis,
Int.
J.
Lab.
Hematol.
42
(2020)
788
e
795
.
[202]
C.
Adamichou,
S.
Georgakis,
G.
Bertsias,
Cytokine
targets
in
lupus
nephritis:
Current
and
future
prospects,
Clin.
Immunol.
206
(2019)
42
e
52
.
[203]
J.H.A.
Creemers,
A.
Ankan,
K.C.B.
Roes,
et
al.,
In
silico
cancer
immunotherapy
trials
uncover
the
consequences
of
therapy-speci
fi
c
response
patterns
for
clinical
trial
design
and
outcome,
Nat.
Commun.
14
(2023),
2348
.
[204]
M.
Zanin,
I.
Chorbev,
B.
Stres,
et
al.,
Community
effort
endorsing
multiscale
modelling,
multiscale
data
science
and
multiscale
computing
for
systems
medicine,
Brief.
Bioinform.
20
(2019)
1057
e
1062
.
[205]
F.T. Musuamba, I. Skottheim Rusten, R. Lesage, et al., Scienti
fi
c and regulatory
evaluation
of
mechanistic
in
silico
drug
and
disease
models
in
drug
devel-
opment:
Building
model
credibility,
CPT
Pharmacometrics
Syst.
Pharmacol.
10
(2021)
804
e
825
.
[206]
A.
Ramella,
F.
Migliavacca,
J.F.
Rodriguez
Matas,
et
al.,
Applicability
assess-
ment for
in-silico
patient-speci
fi
c TEVAR procedures, J. Biomech. 146 (2023),
111423
.
[207]
C.
Montemagno,
J.
Durivault,
C.
Gastaldi,
et
al.,
A
group
of
novel
VEGF
splice
variants
as
alternative
therapeutic
targets
in
renal
cell
carcinoma,
Mol.
Oncol.
17
(2023)
1379
e
1401
.
[208]
A. Giaretta, G. Petrucci, B. Rocca, et al., Physiologically based modelling of the
antiplatelet
effect
of
aspirin:
A
tool
to
characterize
drug
responsiveness
and
inform
precision
dosing,
PLoS
One
17
(2022),
e0268905
.
[209]
A.
Branen,
Y.
Yao,
M.V.
Kothare,
et
al.,
Data
driven
control
of
vagus
nerve
stimulation
for
the
cardiovascular
system:
An
in
silico
computational
study,
Front.
Physiol.
13
(2022),
798157
.
[210]
H.
Zhang,
M.
Yin,
Q.
Liu,
et
al.,
Machine
and
deep
learning-based
clinical
characteristics and laboratory markers for the prediction of sarcopenia, Chin.
Med.
J.
136
(2023)
967
e
973
.
[211]
M. Squires, X. Tao, S. Elangovan, et al., Deep learning and machine learning in
psychiatry:
A
survey
of
current
progress
in
depression
detection,
diagnosis
and
treatment,
Brain
Inform
10
(2023),
10
.
[212]
M. Bertolini, M. Mullen, G. Belitsis, et al., Demonstration of use of a novel 3D
printed
simulator
for
mitral
valve
transcatheter
edge-to-edge
repair
(TEER),
Materials
(Basel)
15
(2022),
4284
.
[213]
L. Swift, C. Zhang, O. Kovalchuk, et al., Dual functionality of the antimicrobial
agent
taurolidine
which
demonstrates
effective
anti-tumor
properties
in
pediatric
neuroblastoma,
Invest,
New
Drugs
38
(2020)
690
e
699
.
[214]
E. Makarev, A.D. Schubert, R.R. Kanherkar, et al.,
In silico
analysis of pathways
activation
landscape
in
oral
squamous
cell
carcinoma
and
oral
leukoplakia,
Cell
Death
Discov.
3
(2017),
17022
.
[215]
V.
Saloura,
E.
Izumchenko,
Z.
Zuo,
et
al.,
Immune
pro
fi
les
in
primary
squa-
mous
cell
carcinoma
of
the
head
and
neck,
Oral
Oncol.
96
(2019)
77
e
88
.
[216]
J.T.
Beck,
M.
Rammage,
G.P.
Jackson,
et
al.,
Arti
fi
cial
intelligence
tool
for
optimizing eligibility screening for clinical trials in a large community cancer
center,
JCO
Clin.
Cancer
Inform
4
(2020)
50
e
59
.
[217]
T.
Haddad,
J.M.
Helgeson,
K.E.
Pomerleau,
et
al.,
Accuracy
of
an
arti
fi
cial
in-
telligence
system
for
cancer
clinical
trial
eligibility
screening:
Retrospective
pilot
study,
JMIR
Med.
Inform
9
(2021),
e27767
.
[218]
L.
Mueller,
P.
Berhanu,
J.
Bouchard,
et
al.,
Application
of
machine
learning
models
to
evaluate
hypoglycemia
risk
in
type
2
diabetes,
Diabetes
Ther.
11
(2020)
681
e
699
.
[219]
M.
Bogart,
Y.
Liu,
T.
Oakland,
et
al.,
Evaluating
triple
therapy
treatment
pathways
in
chronic
obstructive
pulmonary
disease
(COPD):
A
machine-
learning
predictive
model,
Int.
J.
Chron.
Obstruct.
Pulmon.
Dis.
17
(2022)
735
e
747
.
[220]
J.S.
Berger,
L.
Haskell,
W.
Ting,
et
al.,
Evaluation
of
machine
learning
meth-
odology
for
the
prediction
of
healthcare
resource
utilization
and
healthcare
costs in patients with critical limb ischemia
is preventive and personalized
approach
on
the
horizon?
EPMA
J.
11
(2020)
53
e
64
.
[221]
K.G.
Larsen,
J.
Areberg,
D.O.
Åstr
€
om,
Are
self-reported
and
self-monitored
adherence
good
proxies
for
reaching
relevant
plasma
concentrations?:
Ex-
periences from a study of anti-depressants in healthy volunteers, Clin. Trials
18
(2021)
505
e
510
.
[222]
D.L.
Labovitz,
L.
Shafner,
M.
Reyes
Gil,
et
al.,
Using
arti
fi
cial
intelligence
to
reduce
the
risk
of
nonadherence
in
patients
on
anticoagulation
therapy,
Stroke
48
(2017)
1416
e
1419
.
[223]
E.E.
Bain,
L.
Shafner,
D.P.
Walling,
et
al.,
Use
of
a
novel
arti
fi
cial
intelligence
platform on mobile devices to assess dosing compliance in a phase 2 clinical
trial
in
subjects
with
schizophrenia,
JMIR
Mhealth
Uhealth
5
(2017),
e18
.
[224]
A.
Datti,
Academic
drug
discovery
in
an
age
of
research
abundance,
and
the
curious
case
of
chemical
screens
toward
drug
repositioning,
Drug
Discov.
Today
28
(2023),
103522
.
[225]
L.A.B.
P
^
orto,
M.A.F.
Grossi,
E.S.
de
Alecrim,
et
al.,
Deep
venous
thrombosis
in
patients
with
erythema
nodosum
leprosum
in
the
use
of
thalidomide
and
systemic
corticosteroid
in
reference
service
in
Belo
Horizonte,
Minas
Gerais,
Case
Rep.
Dermatol.
Med.
2019
(2019),
8181507
.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
20
[226]
S.
Copley,
P.E.
Yassa,
A.M.
Batterham,
et
al.,
A
clinical
evaluation
of
the
ac-
curacy
of
an
intrathecal
drug
delivery
device,
Neuromodulation
26
(2023)
1240
e
1246
.
[227]
H.
Watanabe,
T.
Hattori,
A.
Kume,
et
al.,
Improved
Parkinsons
disease
motor
score in a single-arm open-label trial of febuxostat and inosine, Medicine 99
(2020),
e21576
.
[228]
D. DeKarske, G. Alva, J.L. Aldred, et al., An open-label, 8-week study of safety
and
ef
fi
cacy
of
pimavanserin
treatment
in
adults
with
Parkinson
’
s
disease
and
depression,
J.
Parkinsons
Dis.
10
(2020)
1751
e
1761
.
[229]
M.
Saber-Ayad,
S.
Hammoudeh,
E.
Abu-Gharbieh,
et
al.,
Current
status
of
baricitinib as a repurposed therapy for COVID-19, Pharmaceuticals (Basel) 14
(2021),
680
.
[230]
P. Richardson, I. Grif
fi
n, C. Tucker, et al., Baricitinib as potential treatment for
2019-nCoV
acute
respiratory
disease,
Lancet
395
(2020)
e30
e
e31
.
[231]
A.C.
Kalil,
J.
Stebbing,
Baricitinib:
The
fi
rst
immunomodulatory
treatment
to
reduce
COVID-19 mortality in a placebo-controlled trial, Lancet Respir. Med.
9
(2021)
1349
e
1351
.
[232]
A.
Kovari,
Explainable
AI
chatbots
towards
XAI
ChatGPT:
A
review,
Heliyon
11
(2025),
e42077
.
[233]
J.M. Brandenburg, B.P. Müller-Stich, M. Wagner, et al., Can surgeons trust AI?
Perspectives
on
machine
learning
in
surgery
and
the
importance
of
eXplainable
Arti
fi
cial
Intelligence
(XAI),
Langenbecks
Arch
Surg.
410
(2025),
53
.
[234]
F.
Wong,
E.J.
Zheng,
J.A.
Valeri,
et
al.,
Discovery
of
a
structural
class
of
anti-
biotics
with
explainable
deep
learning,
Nature
626
(2024)
177
e
185
.
[235]
Q.
Wang,
K.
Huang,
P.
Chandak,
et
al.,
Extending
the
nested
model
for
user-
centric
XAI:
A
design
study
on
GNN-based
drug
repurposing,
IEEE
Trans.
Visual.
Comput.
Graphics
29
(2023)
1266
e
1276
.
[236]
Y.
Gu,
Z.
Yu,
Y.
Wang,
et
al.,
admetSAR3.0:
A
comprehensive
platform
for
exploration,
prediction
and
optimization
of
chemical
ADMET
properties,
Nucleic
Acids
Res.
52
(2024)
W432
e
W438
.
[237]
M.
Picard,
M.
Leclercq,
A.
Bodein,
et
al.,
Improving
drug
repositioning
with
negative
data
labeling
using
large
language
models,
J.
Cheminform.
17
(2025),
16
.
[238]
K.
Raman,
R.
Kumar,
C.J.
Musante,
et
al.,
Integrating
model-informed
drug
development with AI: A synergistic approach to accelerating pharmaceutical
innovation,
Clin.
Transl.
Sci.
18
(2025),
e70124
.
C.
Fu
and
Q.
Chen
Journal
of
Pharmaceutical
Analysis
15
(2025)
101248
21