WEBVTT
Kind: captions
Language: en-US

00:00:02.404 --> 00:00:03.106
[MUSIC] 
 
 

00:00:03.106 --> 00:00:06.800
KATHLEEN SULLIVAN: Welcome&nbsp;
to AI Testing and Evaluation:&nbsp;&nbsp;

00:00:06.800 --> 00:00:11.520
Learnings from Science and Industry.&nbsp;
I’m your host, Kathleen Sullivan. 
 
 

00:00:11.520 --> 00:00:16.640
As generative AI continues to advance, Microsoft&nbsp;
has gathered a range of experts—from genome&nbsp;&nbsp;

00:00:16.640 --> 00:00:21.360
editing to cybersecurity—to share how&nbsp;
their fields approach evaluation and risk&nbsp;&nbsp;

00:00:21.360 --> 00:00:26.000
assessment. Our goal is to learn from&nbsp;
their successes and their stumbles to&nbsp;&nbsp;

00:00:26.000 --> 00:00:31.520
move the science and practice of AI testing&nbsp;
forward. In this series, we’ll explore how&nbsp;&nbsp;

00:00:31.520 --> 00:00:40.304
these insights might help guide the future of AI&nbsp;
development, deployment, and responsible use. 
 
 

00:00:40.293 --> 00:00:40.892
[MUSIC ENDS] 
 
 

00:00:40.892 --> 00:00:45.760
For our final episode of the series, I’m thrilled&nbsp;
to once again be joined by Amanda Craig Deckard,&nbsp;&nbsp;

00:00:45.760 --> 00:00:50.240
senior director of public policy in&nbsp;
Microsoft’s Office of Responsible AI. 
 
 

00:00:50.240 --> 00:00:52.538
Amanda, welcome back to the podcast!  
 
 

00:00:52.538 --> 00:00:53.360
AMANDA CRAIG DECKARD: Thank you so much. 
 
 

00:00:53.360 --> 00:00:57.840
SULLIVAN: In our intro episode, you really&nbsp;
helped set the stage for this series. And&nbsp;&nbsp;

00:00:57.840 --> 00:01:03.280
it’s been great, because since then, we’ve had&nbsp;
the pleasure of speaking with governance experts&nbsp;&nbsp;

00:01:03.280 --> 00:01:10.400
about genome editing, pharma, medical devices,&nbsp;
cybersecurity, and we’ve also gotten to spend&nbsp;&nbsp;

00:01:10.400 --> 00:01:15.840
some time with our own Microsoft responsible&nbsp;
AI leaders and hear reflections from them.  
 
 

00:01:15.840 --> 00:01:20.000
And here’s what stuck with me, and I’d&nbsp;
love to hear from you on this, as well:&nbsp;&nbsp;

00:01:20.000 --> 00:01:26.800
testing builds trust; context is shaping risk;&nbsp;
and every field is really thinking about striking&nbsp;&nbsp;

00:01:26.800 --> 00:01:32.080
its own balance between pre-deployment&nbsp;
testing and post-deployment monitoring.  
 
 

00:01:32.080 --> 00:01:37.280
So drawing on what you’ve learned from&nbsp;
the workshop and the case studies,&nbsp;&nbsp;

00:01:37.280 --> 00:01:42.625
what headline insights do you think&nbsp;
matter the most for AI governance? 
 
 

00:01:42.625 --> 00:01:46.240
CRAIG DECKARD: It's been really interesting&nbsp;
to learn from all of these different domains,&nbsp;&nbsp;

00:01:46.240 --> 00:01:50.560
and there are, you know, lots of&nbsp;
really interesting takeaways.  
 
 

00:01:50.560 --> 00:01:56.800
I think a starting point for me is actually pretty&nbsp;
similar to where you landed, which is just that&nbsp;&nbsp;

00:01:56.800 --> 00:02:05.040
testing is really important for trust, and it's&nbsp;
also really hard [LAUGHS] to figure out exactly,&nbsp;&nbsp;

00:02:05.040 --> 00:02:09.280
you know, how to get it right, how to&nbsp;
make sure that you're addressing risks,&nbsp;&nbsp;

00:02:09.280 --> 00:02:16.560
that you're not constraining innovation, that you&nbsp;
are recognizing that a lot of the industry that's&nbsp;&nbsp;

00:02:16.560 --> 00:02:20.960
impacted is really different. You have small&nbsp;
organizations, you have large organizations,&nbsp;&nbsp;

00:02:20.960 --> 00:02:25.440
and you want to enable that opportunity that is&nbsp;
enabled by the technology across the board.  
 
 

00:02:25.440 --> 00:02:30.320
And so it's just difficult to, kind&nbsp;
of, get all of these dynamics right,&nbsp;&nbsp;

00:02:30.320 --> 00:02:34.960
especially when, you know, I think we heard&nbsp;
from other domains, testing is not some,&nbsp;&nbsp;

00:02:34.960 --> 00:02:38.960
sort of, like, oh, simple thing, right.&nbsp;
There's not this linear path from, like,&nbsp;&nbsp;

00:02:38.960 --> 00:02:41.360
A to B where you just test the&nbsp;
one thing and you're done.  
 
 

00:02:41.360 --> 00:02:41.985
SULLIVAN: Right. 
 
 

00:02:41.985 --> 00:02:47.280
CRAIG DECKARD: It's complex, right. Testing&nbsp;
is multistage. There's a lot of testing by&nbsp;&nbsp;

00:02:47.280 --> 00:02:53.680
different actors. There are a lot of different&nbsp;
purposes for which you might test. As I think&nbsp;&nbsp;

00:02:53.680 --> 00:02:58.640
it was Dan Carpenter who talked about it's not&nbsp;
just about testing for safety. It's also about&nbsp;&nbsp;

00:02:58.640 --> 00:03:04.640
testing for efficacy and building confidence&nbsp;
in the right dosage for pharmaceuticals,&nbsp;&nbsp;

00:03:04.640 --> 00:03:07.920
for example. And that's across the&nbsp;
board for all of these domains,&nbsp;&nbsp;

00:03:07.920 --> 00:03:12.080
right. That you're really thinking about&nbsp;
the performance of the technology. You're&nbsp;&nbsp;

00:03:12.080 --> 00:03:16.960
thinking about safety. You're trying&nbsp;
to also calibrate for efficiency.  
 
 

00:03:16.960 --> 00:03:23.120
And so those tradeoffs, every expert&nbsp;
shared that navigating those is really&nbsp;&nbsp;

00:03:23.120 --> 00:03:29.920
challenging. And also that there were real&nbsp;
impacts to early choices in the, sort of,&nbsp;&nbsp;

00:03:29.920 --> 00:03:34.880
governance of risk in these different domains&nbsp;
and the development of the testing, sort of,&nbsp;&nbsp;

00:03:34.880 --> 00:03:38.160
expectations, and that in some cases,&nbsp;
this had been difficult to reverse,&nbsp;&nbsp;

00:03:38.160 --> 00:03:42.480
which also just layers on that complexity&nbsp;
and that difficulty in a different way.&nbsp;&nbsp;

00:03:42.480 --> 00:03:47.760
So that’s the super high-level takeaway. But&nbsp;
maybe if I could just quickly distill, like,&nbsp;&nbsp;

00:03:47.760 --> 00:03:54.080
three takeaways that I think really are applicable&nbsp;
to AI in a bit more of a granular way.  
 
 

00:03:54.080 --> 00:04:02.080
You know, one is about, how is the testing&nbsp;
exactly used? For what purpose? And the&nbsp;&nbsp;

00:04:02.080 --> 00:04:07.520
second is what emphasis there is on this&nbsp;
pre- versus post-deployment testing and&nbsp;&nbsp;

00:04:07.520 --> 00:04:13.600
monitoring. And then the third is how&nbsp;
rigid versus adaptive the, sort of,&nbsp;&nbsp;

00:04:13.600 --> 00:04:17.520
testing regimes or frameworks are&nbsp;
in these different domains.  
 
 

00:04:17.520 --> 00:04:24.720
So on the first—how is testing used?—so is&nbsp;
testing something that impacts market entry,&nbsp;&nbsp;

00:04:24.720 --> 00:04:32.160
for example? Or is it something that might be used&nbsp;
more for informing how risk is evolving in the&nbsp;&nbsp;

00:04:32.160 --> 00:04:38.720
domain and how broader risk management strategies&nbsp;
might need to be applied? We have examples,&nbsp;&nbsp;

00:04:38.720 --> 00:04:43.920
like the pharmaceutical or medical device industry&nbsp;
experts with whom you spoke, that's really,&nbsp;&nbsp;

00:04:43.920 --> 00:04:49.280
you know, testing … there is a pre-deployment&nbsp;
requirement. So that's one question.  
 
 

00:04:49.280 --> 00:04:56.560
The second is this emphasis on pre- versus&nbsp;
post-deployment testing and monitoring,&nbsp;&nbsp;

00:04:56.560 --> 00:05:02.400
and we really did see across domains&nbsp;
that in many cases, there is a desire&nbsp;&nbsp;

00:05:02.400 --> 00:05:06.160
for both pre- and post-deployment,&nbsp;
sort of, testing and monitoring,&nbsp;&nbsp;

00:05:06.160 --> 00:05:12.560
but also that, sort of, naturally in these&nbsp;
different domains, a degree of emphasis on&nbsp;&nbsp;

00:05:12.560 --> 00:05:19.040
one or the other had evolved and that had a&nbsp;
real impact on governance and tradeoffs.  
 
 

00:05:19.040 --> 00:05:24.240
And the third is just how rigid versus&nbsp;
adaptive these testing and evaluation&nbsp;&nbsp;

00:05:24.240 --> 00:05:29.200
regimes or frameworks are in these different&nbsp;
domains. We saw, you know, in some domains,&nbsp;&nbsp;

00:05:29.200 --> 00:05:37.200
the testing requirements were more rigid as&nbsp;
you might expect in more of the pharmaceutical&nbsp;&nbsp;

00:05:37.200 --> 00:05:43.440
or medical devices industries, for example. And&nbsp;
in other domains, there was this more, sort of,&nbsp;&nbsp;

00:05:43.440 --> 00:05:50.480
adaptive approach to how testing might get&nbsp;
used. So, for example, in the case of our&nbsp;&nbsp;

00:05:50.480 --> 00:05:56.240
other general-purpose technologies, you know,&nbsp;
you spoke with Alta Charo on genome editing,&nbsp;&nbsp;

00:05:56.240 --> 00:06:01.920
and in our case studies, we also explored&nbsp;
this in the context of nanotechnology. In&nbsp;&nbsp;

00:06:01.920 --> 00:06:09.840
those general-purpose technology domains, there is&nbsp;
more emphasis on downstream or application-context&nbsp;&nbsp;

00:06:09.840 --> 00:06:17.680
testing that is more, sort of, adaptive to the use&nbsp;
scenario of the technology and, you know, having&nbsp;&nbsp;

00:06:17.680 --> 00:06:23.200
that work in conjunction with testing more at&nbsp;
the, kind of, level of the technology itself. 
 
 

00:06:23.200 --> 00:06:26.560
SULLIVAN: I want to double-click on&nbsp;
a number of the things we just talked&nbsp;&nbsp;

00:06:26.560 --> 00:06:29.840
about. But actually, before we go too much deeper,&nbsp;&nbsp;

00:06:29.840 --> 00:06:34.000
a question on if there's anything that really&nbsp;
surprised you or challenged maybe some of your&nbsp;&nbsp;

00:06:34.000 --> 00:06:39.665
own assumptions in this space from some of the&nbsp;
discussions that we had over the series. 
 
 

00:06:39.665 --> 00:06:43.920
CRAIG DECKARD: Yeah. You know, I know I've already&nbsp;
just mentioned this pre- versus post-deployment&nbsp;&nbsp;

00:06:43.920 --> 00:06:50.160
testing and monitoring issue, but it was&nbsp;
something that was very interesting to me and&nbsp;&nbsp;

00:06:50.160 --> 00:06:56.960
in some ways surprised me or made me just realize&nbsp;
something that I hadn't fully connected before,&nbsp;&nbsp;

00:06:56.960 --> 00:07:03.440
about how these, sort of, regimes might evolve&nbsp;
in different contexts and why. And in part,&nbsp;&nbsp;

00:07:03.440 --> 00:07:10.240
I couldn't help but bring the context I have&nbsp;
from cybersecurity policy into this, kind of,&nbsp;&nbsp;

00:07:10.240 --> 00:07:14.560
processing of what we learned and&nbsp;
reflection because there was a real&nbsp;&nbsp;

00:07:14.560 --> 00:07:20.080
contrast for me between the pharmaceutical&nbsp;
industry and the cybersecurity domain when&nbsp;&nbsp;

00:07:20.080 --> 00:07:24.240
I think about the emphasis on pre-&nbsp;
versus post-deployment monitoring. 
 
 

00:07:24.240 --> 00:07:31.680
And on the one hand, we have in the pharmaceutical&nbsp;
domain a real emphasis that has developed around&nbsp;&nbsp;

00:07:31.680 --> 00:07:38.160
pre-market testing. And there is also an&nbsp;
expectation in some circumstances in the&nbsp;&nbsp;

00:07:38.160 --> 00:07:43.280
pharmaceutical domain for post-deployment&nbsp;
testing, as well. But as we learned from&nbsp;&nbsp;

00:07:43.280 --> 00:07:51.280
our experts in that domain, there has naturally&nbsp;
been a real, kind of, emphasis on the pre-market&nbsp;&nbsp;

00:07:51.280 --> 00:07:58.240
portion of that testing. And in reality, even&nbsp;
where post-market monitoring is required and&nbsp;&nbsp;

00:07:58.240 --> 00:08:04.400
post-market testing is required, it does not&nbsp;
always actually happen. And the experts really&nbsp;&nbsp;

00:08:04.400 --> 00:08:09.440
explained that, you know, part of it is just the&nbsp;
incentive structure around the emphasis around,&nbsp;&nbsp;

00:08:09.440 --> 00:08:16.160
you know, the testing as a pre-market, sort of,&nbsp;
entry requirement. And also just the resources&nbsp;&nbsp;

00:08:16.160 --> 00:08:22.400
that exist among regulators, right. There's&nbsp;
limited resources, right. And so there are just&nbsp;&nbsp;

00:08:22.400 --> 00:08:26.720
choices and tradeoffs that they need to make&nbsp;
in their own, sort of, enforcement work.  
 
 

00:08:26.720 --> 00:08:34.320
And then on the other hand, you know, in&nbsp;
cybersecurity, I never thought about the, kind of,&nbsp;&nbsp;

00:08:34.320 --> 00:08:40.560
emphasis on things like coordinated vulnerability&nbsp;
disclosure and bug bounties that have really&nbsp;&nbsp;

00:08:40.560 --> 00:08:47.280
developed in the cybersecurity domain. But it's a&nbsp;
really important part of how we secure technology&nbsp;&nbsp;

00:08:47.280 --> 00:08:53.600
and enhance cybersecurity over time, where we have&nbsp;
these norms that have developed where, you know,&nbsp;&nbsp;

00:08:53.600 --> 00:08:58.960
security researchers are doing really important&nbsp;
research. They're finding vulnerabilities in&nbsp;&nbsp;

00:08:58.960 --> 00:09:03.920
products. And we have norms developed where&nbsp;
they report those to the companies that are in a&nbsp;&nbsp;

00:09:03.920 --> 00:09:09.280
position to address those vulnerabilities. And in&nbsp;
some cases, those companies actually pay, through&nbsp;&nbsp;

00:09:09.280 --> 00:09:15.520
bug bounties, the researchers. And perhaps in&nbsp;
some ways, the role of coordinated vulnerability&nbsp;&nbsp;

00:09:15.520 --> 00:09:20.400
disclosure and bug bounties has evolved the way&nbsp;
that it has because there hasn't been as much&nbsp;&nbsp;

00:09:20.400 --> 00:09:26.720
emphasis on the pre-market testing across the&nbsp;
board at least in the context of software.  
 
 

00:09:26.720 --> 00:09:32.080
And so you look at those two industries and it&nbsp;
was interesting to me to study them to some extent&nbsp;&nbsp;

00:09:32.080 --> 00:09:39.200
in contrast with each other as this way that&nbsp;
the incentives and the resources that need to&nbsp;&nbsp;

00:09:39.200 --> 00:09:45.680
be applied to testing, sort of, evolve to address&nbsp;
where there's, kind of, more or less emphasis. 
 
 

00:09:45.680 --> 00:09:49.120
SULLIVAN: It's a great point. I mean, I&nbsp;
think what we're hearing—and what you're&nbsp;&nbsp;

00:09:49.120 --> 00:09:52.400
saying—is just exactly this choice …&nbsp;
like, is there a binary choice between&nbsp;&nbsp;

00:09:52.400 --> 00:09:57.120
focusing on pre-deployment testing or&nbsp;
post-deployment monitoring? And, you know,&nbsp;&nbsp;

00:09:57.120 --> 00:10:01.520
I think our assumption is that we need to do&nbsp;
both. But I'd love to hear from you on that. 
 
 

00:10:01.520 --> 00:10:08.320
CRAIG DECKARD: Absolutely. I think we need to&nbsp;
do both. I'm very persuaded by this inclination&nbsp;&nbsp;

00:10:08.320 --> 00:10:14.560
always that there's value in trying to really&nbsp;
do it all in a risk management context.  
 
 

00:10:14.560 --> 00:10:20.640
And also, we know one of the principles of risk&nbsp;
management is you have to prioritize because there&nbsp;&nbsp;

00:10:20.640 --> 00:10:27.120
are finite resources. And I think that's where we&nbsp;
get to this challenge in really thinking deeply,&nbsp;&nbsp;

00:10:27.120 --> 00:10:32.240
especially as we're in the early days of AI&nbsp;
governance, and we need to be very thoughtful&nbsp;&nbsp;

00:10:32.240 --> 00:10:37.680
about, you know, tradeoffs that we may not&nbsp;
want to be making but we are because, again,&nbsp;&nbsp;

00:10:37.680 --> 00:10:42.480
these are finite choices and we, kind of,&nbsp;
can't help but put our finger on the dial in&nbsp;&nbsp;

00:10:42.480 --> 00:10:47.280
different directions with our choices that, you&nbsp;
know, it's going to be very difficult to have,&nbsp;&nbsp;

00:10:47.280 --> 00:10:52.000
sort of, equal emphasis on both. And we&nbsp;
need to invest in both, but we need to be&nbsp;&nbsp;

00:10:52.000 --> 00:10:59.840
very deliberate about the roles of each and how&nbsp;
they complement each other and who does which&nbsp;&nbsp;

00:10:59.840 --> 00:11:05.280
and how we use what we learn from pre- versus&nbsp;
post-deployment testing and monitoring. 
 
 

00:11:05.280 --> 00:11:09.280
SULLIVAN: Maybe just spending a little bit more&nbsp;
time here … you know, a lot of attention goes&nbsp;&nbsp;

00:11:09.280 --> 00:11:14.400
into testing models upstream, but risk often&nbsp;
shows up once they're wired into real products&nbsp;&nbsp;

00:11:14.400 --> 00:11:20.400
and workflows. How much does deployment context&nbsp;
change the risk picture from your perspective? 
 
 

00:11:20.400 --> 00:11:24.640
CRAIG DECKARD: Yeah, I … such an&nbsp;
important question. I really agree that&nbsp;&nbsp;

00:11:24.640 --> 00:11:29.920
there has been a lot of emphasis to date&nbsp;
on, sort of, testing models upstream,&nbsp;&nbsp;

00:11:29.920 --> 00:11:39.280
the AI model evaluation. And it's also really&nbsp;
important that we bring more attention into&nbsp;&nbsp;

00:11:39.280 --> 00:11:46.960
evaluation at the system or application level. And&nbsp;
I actually see that in governance conversations,&nbsp;&nbsp;

00:11:46.960 --> 00:11:54.640
this is actually increasingly raised, this need&nbsp;
to have system-level evaluation. We see this&nbsp;&nbsp;

00:11:54.640 --> 00:12:00.480
across regulation. We also see it in the context&nbsp;
of just organizations trying to put in governance&nbsp;&nbsp;

00:12:00.480 --> 00:12:04.640
requirements for how their organization is going&nbsp;
to operate in deploying this technology.  
 
 

00:12:04.640 --> 00:12:11.040
And there's a gap today in terms of best&nbsp;
practices around system-level testing,&nbsp;&nbsp;

00:12:11.040 --> 00:12:19.280
perhaps even more than model-level evaluation. And&nbsp;
it's really important because in a lot of cases,&nbsp;&nbsp;

00:12:19.280 --> 00:12:25.680
the deployment context really does impact&nbsp;
the risk picture, especially with AI,&nbsp;&nbsp;

00:12:25.680 --> 00:12:30.560
which is a general-purpose technology,&nbsp;
and we really saw this in our study of&nbsp;&nbsp;

00:12:30.560 --> 00:12:34.320
other domains that represented&nbsp;
general-purpose technology.  
 
 

00:12:34.320 --> 00:12:40.640
So in the case study that you can find online&nbsp;
on nanotechnology, you know, there's a real&nbsp;&nbsp;

00:12:40.640 --> 00:12:48.640
distinction between the risk evaluation and&nbsp;
the governance of nanotechnology in different&nbsp;&nbsp;

00:12:48.640 --> 00:12:55.840
deployment contexts. So the chapter that our&nbsp;
expert on nanotechnology wrote really goes into&nbsp;&nbsp;

00:12:55.840 --> 00:13:01.200
incredibly interesting detail around, you know,&nbsp;
deployment of nanotechnology in the context of,&nbsp;&nbsp;

00:13:01.200 --> 00:13:07.440
like, chemical applications versus consumer&nbsp;
electronics versus pharmaceuticals versus&nbsp;&nbsp;

00:13:07.440 --> 00:13:13.280
construction and how the way that nanoparticles&nbsp;
are basically delivered in all those different&nbsp;&nbsp;

00:13:13.280 --> 00:13:19.520
deployment contexts, as well as, like, what&nbsp;
the risk of the actual use scenario is just&nbsp;&nbsp;

00:13:19.520 --> 00:13:25.760
varies so much. And so there's a real need to&nbsp;
do that kind of risk evaluation and testing in&nbsp;&nbsp;

00:13:25.760 --> 00:13:30.160
the deployment context, and this difference&nbsp;
in terms of risks and what we learned in&nbsp;&nbsp;

00:13:30.160 --> 00:13:36.720
these other domains where, you know, there are&nbsp;
these different approaches to trying to really&nbsp;&nbsp;

00:13:36.720 --> 00:13:41.600
think about and gain efficiencies and address&nbsp;
risks at a horizontal level versus, you know,&nbsp;&nbsp;

00:13:41.600 --> 00:13:47.760
taking a real sector-by-sector approach. And&nbsp;
to some extent, it seems like it's more time&nbsp;&nbsp;

00:13:47.760 --> 00:13:54.080
intensive to do that sectoral deployment-specific&nbsp;
work. And at the same time, perhaps there are&nbsp;&nbsp;

00:13:54.080 --> 00:14:00.320
efficiencies to be gained by actually doing&nbsp;
the work in the context in which, you know, you&nbsp;&nbsp;

00:14:00.320 --> 00:14:07.040
have a better understanding of the risk that can&nbsp;
result from really deploying this technology.  
 
 

00:14:07.040 --> 00:14:12.560
And ultimately, [LAUGHS] really what we also&nbsp;
need to think about here is probably, in the end,&nbsp;&nbsp;

00:14:12.560 --> 00:14:17.200
just like pre- and post-deployment testing,&nbsp;
you need both. Not probably; certainly!  
 
 

00:14:17.200 --> 00:14:24.080
So effectively we need to think about evaluation&nbsp;
at the model level and the system level as being&nbsp;&nbsp;

00:14:24.080 --> 00:14:29.360
really important. And it's really important&nbsp;
to get system evaluation right so that we can&nbsp;&nbsp;

00:14:29.360 --> 00:14:36.160
actually get trust in this technology in&nbsp;
deployment context so we enable adoption&nbsp;&nbsp;

00:14:37.040 --> 00:14:43.840
in low- and in high-risk deployments in a way that&nbsp;
means that we've done risk evaluation in each of&nbsp;&nbsp;

00:14:43.840 --> 00:14:49.520
those contexts in a way that really makes sense in&nbsp;
terms of the resources that we need to apply and&nbsp;&nbsp;

00:14:49.520 --> 00:14:56.760
ultimately we are able to unlock more applications&nbsp;
of this technology in a risk-informed way. 
 
 

00:14:56.760 --> 00:15:01.840
SULLIVAN: That's great. I mean, I couldn't agree&nbsp;
more. I think these contexts, the approaches are&nbsp;&nbsp;

00:15:01.840 --> 00:15:07.600
so important for trust and adoption, and I'd love&nbsp;
to hear from you, what do we need to advance AI&nbsp;&nbsp;

00:15:07.600 --> 00:15:12.800
evaluation and testing in our ecosystem? What&nbsp;
are some of the big gaps that you're seeing,&nbsp;&nbsp;

00:15:12.800 --> 00:15:18.480
and what role can different stakeholders play&nbsp;
in filling them? And maybe an add-on, actually:&nbsp;&nbsp;

00:15:18.480 --> 00:15:23.505
is there some sort of network effect&nbsp;
that could 10x our testing capacity? 
 
 

00:15:23.505 --> 00:15:29.280
CRAIG DECKARD: Absolutely. So there's&nbsp;
a lot of work that needs to be done,&nbsp;&nbsp;

00:15:29.280 --> 00:15:35.280
and there's a lot of work in process to&nbsp;
really level up our whole evaluation and&nbsp;&nbsp;

00:15:35.280 --> 00:15:41.760
testing ecosystem. We learned, across&nbsp;
domains, that there’s really a need to&nbsp;&nbsp;

00:15:41.760 --> 00:15:48.720
advance our thinking and our practice&nbsp;
in three areas: rigor of testing;&nbsp;&nbsp;

00:15:48.720 --> 00:15:55.200
standardization of methodologies and processes;&nbsp;
and interpretability of test results.  
 
 

00:15:55.200 --> 00:16:01.280
So what we mean by rigor is that&nbsp;
we are ensuring that what we are&nbsp;&nbsp;

00:16:01.280 --> 00:16:05.920
ultimately evaluating in terms of risks&nbsp;
is defined in a scientifically valid way&nbsp;&nbsp;

00:16:05.920 --> 00:16:12.320
and we are able to measure against that&nbsp;
risk in a scientifically valid way.  
 
 

00:16:12.320 --> 00:16:19.200
By standardization, what we mean is that there's&nbsp;
really an accepted and well-understood and,&nbsp;&nbsp;

00:16:19.200 --> 00:16:25.280
again, a scientifically valid&nbsp;
methodology for doing that testing&nbsp;&nbsp;

00:16:25.280 --> 00:16:33.200
and for actually producing artifacts out&nbsp;
of that testing that are meeting those&nbsp;&nbsp;

00:16:33.200 --> 00:16:39.760
standards. And that sets us up for the final&nbsp;
portion on interpretability, which is, like,&nbsp;&nbsp;

00:16:39.760 --> 00:16:45.520
really the process by which you can trust that&nbsp;
the testing has been done in this rigorous and&nbsp;&nbsp;

00:16:45.520 --> 00:16:52.560
standardized way and that then you have artifacts&nbsp;
that result from the testing process that can&nbsp;&nbsp;

00:16:52.560 --> 00:16:58.160
really be used in the risk management context&nbsp;
because they can be interpreted, right.  
 
 

00:16:58.160 --> 00:17:05.760
We understand how to, like, apply weight to them&nbsp;
for our risk-management decisions. We actually are&nbsp;&nbsp;

00:17:05.760 --> 00:17:11.680
able to interpret them in a way that perhaps&nbsp;
they inform other downstream risk mitigations&nbsp;&nbsp;

00:17:11.680 --> 00:17:16.320
that address the risks that we see through the&nbsp;
testing results and that we actually understand&nbsp;&nbsp;

00:17:16.320 --> 00:17:23.040
what limitations apply to the test results and&nbsp;
why they may or may not be valid in certain,&nbsp;&nbsp;

00:17:23.040 --> 00:17:27.680
sort of, deployment contexts, for example,&nbsp;
and especially in the context of other risk&nbsp;&nbsp;

00:17:27.680 --> 00:17:33.760
mitigations that we need to apply. So there's a&nbsp;
need to advance all three of those things—rigor,&nbsp;&nbsp;

00:17:33.760 --> 00:17:40.640
standardization, and interpretability—to level up&nbsp;
the whole testing and evaluation ecosystem.  
 
 

00:17:40.640 --> 00:17:48.480
And when we think about what actors should&nbsp;
be involved in that work … really everybody,&nbsp;&nbsp;

00:17:48.480 --> 00:17:54.400
which is both complex to orchestrate but&nbsp;
also really important. And so, you know,&nbsp;&nbsp;

00:17:54.400 --> 00:18:01.040
you need to have the entire value chain&nbsp;
involved in really advancing this work.&nbsp;&nbsp;

00:18:01.040 --> 00:18:05.680
You need the model developers, but you also&nbsp;
need the system developers and deployers that&nbsp;&nbsp;

00:18:05.680 --> 00:18:12.480
are really engaged in advancing the science&nbsp;
of evaluation and advancing how we are using&nbsp;&nbsp;

00:18:12.480 --> 00:18:16.880
these testing artifacts in the&nbsp;
risk management process.  
 
 

00:18:16.880 --> 00:18:24.080
When we think about what could actually 10x our&nbsp;
testing capacity—that's the dream, right? We all&nbsp;&nbsp;

00:18:24.080 --> 00:18:29.440
want to accelerate our progress in this space.&nbsp;
You know, I think we need work across all three&nbsp;&nbsp;

00:18:29.440 --> 00:18:34.560
of those areas of rigor, standardization, and&nbsp;
interpretability, but I think one that will really&nbsp;&nbsp;

00:18:34.560 --> 00:18:43.120
help accelerate our progress across the board is&nbsp;
that standardization work, because ultimately,&nbsp;&nbsp;

00:18:43.120 --> 00:18:50.240
you're going to need to have these tests be done&nbsp;
and applied across so many different contexts,&nbsp;&nbsp;

00:18:50.240 --> 00:18:54.880
and ultimately, while we want the whole&nbsp;
value chain engaged in the development&nbsp;&nbsp;

00:18:54.880 --> 00:18:59.280
of the thinking and the science and the&nbsp;
standards in this space, we also need&nbsp;&nbsp;

00:18:59.280 --> 00:19:05.360
to realize that not every organization is&nbsp;
necessarily going to have the capacity to,&nbsp;&nbsp;

00:19:05.360 --> 00:19:10.880
kind of, contribute to developing the ways&nbsp;
that we create and use these tests. And&nbsp;&nbsp;

00:19:10.880 --> 00:19:14.080
there are going to be many organizations&nbsp;
that are going to benefit from there being&nbsp;&nbsp;

00:19:14.080 --> 00:19:20.640
standardization of the methodologies and the&nbsp;
artifacts that they can pick up and use.  
 
 

00:19:20.640 --> 00:19:26.320
One thing that I know we've heard throughout this&nbsp;
podcast series from our experts in other domains,&nbsp;&nbsp;

00:19:26.320 --> 00:19:31.360
including Timo [Minssen] in the medical&nbsp;
devices context and Ciaran [Martin] in&nbsp;&nbsp;

00:19:31.360 --> 00:19:37.040
the cybersecurity context, is that there's been&nbsp;
a recognition, as those domains have evolved,&nbsp;&nbsp;

00:19:37.040 --> 00:19:42.320
that there's a need to calibrate our, sort of,&nbsp;
expectations for different actors in the ecosystem&nbsp;&nbsp;

00:19:42.320 --> 00:19:46.320
and really understand that small businesses,&nbsp;
for example, just cannot apply the same degree&nbsp;&nbsp;

00:19:46.320 --> 00:19:52.720
of resources that others may be able to, to do&nbsp;
testing and evaluation and risk management. And&nbsp;&nbsp;

00:19:52.720 --> 00:19:59.760
so the benefit of having standardized approaches&nbsp;
is that those organizations are able to, kind of,&nbsp;&nbsp;

00:19:59.760 --> 00:20:03.600
integrate into the broader supply chain&nbsp;
ecosystem and apply their own, kind of,&nbsp;&nbsp;

00:20:03.600 --> 00:20:09.520
risk management practices in their own&nbsp;
context in a way that is more efficient.  
 
 

00:20:09.520 --> 00:20:14.160
And finally, the last stakeholder that I&nbsp;
think is really important to think about&nbsp;&nbsp;

00:20:14.160 --> 00:20:18.400
in terms of partnership across the&nbsp;
ecosystem to really advance the whole&nbsp;&nbsp;

00:20:18.400 --> 00:20:23.680
testing and evaluation work that needs&nbsp;
to happen is government partners, right,&nbsp;&nbsp;

00:20:23.680 --> 00:20:28.800
and thinking beyond the value chain, the AI supply&nbsp;
chain, and really thinking about public-private&nbsp;&nbsp;

00:20:28.800 --> 00:20:35.360
partnership. That's going to be incredibly&nbsp;
important to advancing this ecosystem.  
 
 

00:20:35.360 --> 00:20:42.800
You know, I think there's been real progress&nbsp;
already in the AI evaluation and testing&nbsp;&nbsp;

00:20:42.800 --> 00:20:49.600
ecosystem in the public-private partnership&nbsp;
context. We have been really supportive of&nbsp;&nbsp;

00:20:49.600 --> 00:20:55.440
the work of the International Network of AI&nbsp;
Safety and Security Institutes and the Center&nbsp;&nbsp;

00:20:55.440 --> 00:21:02.320
for AI Standards and Innovation that all allow&nbsp;
for that kind of public-private partnership on&nbsp;&nbsp;

00:21:02.320 --> 00:21:09.520
actually testing and advancing the science&nbsp;
and best practices around standards. 
 
 

00:21:09.520 --> 00:21:14.800
And there are other innovative, kind of,&nbsp;
partnerships, as well, in the ecosystem. You know,&nbsp;&nbsp;

00:21:14.800 --> 00:21:20.880
Singapore has recently launched their Global AI&nbsp;
Assurance Pilot findings. And that effort really&nbsp;&nbsp;

00:21:20.880 --> 00:21:27.280
paired application deployers and testers so that&nbsp;
consequential impacts at deployment could really&nbsp;&nbsp;

00:21:27.280 --> 00:21:34.480
be tested. And that's a really fruitful, sort&nbsp;
of, effort that complements the work of these&nbsp;&nbsp;

00:21:34.480 --> 00:21:38.960
institutes and centers that are more focused on&nbsp;
evaluation at the model level, for example.  
 
 

00:21:39.600 --> 00:21:44.480
And in general, you know, I think that there's&nbsp;
just really a lot of benefits for us thinking&nbsp;&nbsp;

00:21:44.480 --> 00:21:49.600
expansively about what we can accomplish through&nbsp;
deep, meaningful public-private partnership in&nbsp;&nbsp;

00:21:49.600 --> 00:21:55.120
this space. I'm really excited to see where we&nbsp;
can go from here with building on, you know,&nbsp;&nbsp;

00:21:55.120 --> 00:22:01.280
partnerships across AI supply chains and with&nbsp;
governments and public-private partnerships.  
 
 

00:22:01.280 --> 00:22:05.280
SULLIVAN: I couldn't agree more. I mean, this&nbsp;
notion of more engagement across the ecosystem and&nbsp;&nbsp;

00:22:05.280 --> 00:22:11.280
value chain is super important for us and informs&nbsp;
how we think about the space completely.  
 
 

00:22:11.280 --> 00:22:15.680
If you could invite any other&nbsp;
industry to the next workshop,&nbsp;&nbsp;

00:22:15.680 --> 00:22:18.640
maybe quantum safety, space tech, even gaming,&nbsp;&nbsp;

00:22:18.640 --> 00:22:22.865
who's on your wish list? And maybe what are some&nbsp;
of the things you'd want to go deeper on? 
 
 

00:22:22.865 --> 00:22:29.680
CRAIG DECKARD: This is something that we really&nbsp;
welcome feedback on if anyone listening has ideas&nbsp;&nbsp;

00:22:29.680 --> 00:22:33.440
about other domains that would be interesting&nbsp;
to study. I will say, I think I shared at the&nbsp;&nbsp;

00:22:33.440 --> 00:22:40.880
outset of this podcast series, the domains that&nbsp;
we added in this round of our efforts in studying&nbsp;&nbsp;

00:22:40.880 --> 00:22:46.320
other domains actually all came from feedback&nbsp;
that we received from, you know, folks we’d&nbsp;&nbsp;

00:22:46.320 --> 00:22:51.920
engaged with our first study of other domains and&nbsp;
multilateral, sort of, governance institutions.&nbsp;&nbsp;

00:22:51.920 --> 00:22:57.440
And so we're really keen to think about what&nbsp;
other domains could be interesting to study. And&nbsp;&nbsp;

00:22:57.440 --> 00:23:05.040
we are also keen to go deeper, building on what we&nbsp;
learned in this round of effort going forward.  
 
 

00:23:05.040 --> 00:23:10.560
One of the areas that I am particularly&nbsp;
really interested in is going deeper on,&nbsp;&nbsp;

00:23:10.560 --> 00:23:17.280
what, sort of, transparency and information&nbsp;
sharing about risk evaluation and testing&nbsp;&nbsp;

00:23:17.280 --> 00:23:24.000
will be really useful to share in different&nbsp;
contexts? So across the AI supply chain,&nbsp;&nbsp;

00:23:24.000 --> 00:23:30.640
what is the information that's going to be&nbsp;
really meaningful to share between developers&nbsp;&nbsp;

00:23:30.640 --> 00:23:35.440
and deployers of models and systems and those&nbsp;
that are ultimately using this technology in&nbsp;&nbsp;

00:23:35.440 --> 00:23:42.880
particular deployment contexts? And, you know,&nbsp;
I think that we could have much to learn from&nbsp;&nbsp;

00:23:42.880 --> 00:23:50.800
other general-purpose technologies like genome&nbsp;
editing and nanotechnology and cybersecurity,&nbsp;&nbsp;

00:23:50.800 --> 00:23:57.120
where we could learn a bit more about the kinds&nbsp;
of information that they have shared across the&nbsp;&nbsp;

00:23:57.120 --> 00:24:02.960
development and deployment life cycle and&nbsp;
how that has strengthened risk management&nbsp;&nbsp;

00:24:02.960 --> 00:24:07.680
in general as well as provided a really strong&nbsp;
feedback loop around testing and evaluation.&nbsp;&nbsp;

00:24:08.560 --> 00:24:13.200
What kind of testing is most useful&nbsp;
to do at what point in the life cycle,&nbsp;&nbsp;

00:24:13.200 --> 00:24:19.360
and what artifacts are most useful to share as a&nbsp;
result of that testing and evaluation work?  
 
 

00:24:19.360 --> 00:24:26.240
I'll say, as Microsoft, we have been really&nbsp;
investing in how we are sharing information&nbsp;&nbsp;

00:24:26.240 --> 00:24:32.400
with our various stakeholders. We also have been&nbsp;
engaged with others in industry in reporting what&nbsp;&nbsp;

00:24:32.400 --> 00:24:38.800
we've done in the context of the Hiroshima AI&nbsp;
Process, or HAIP, Reporting Framework. This is an&nbsp;&nbsp;

00:24:38.800 --> 00:24:48.320
effort that is really just in its first round of&nbsp;
really exploring how this kind of reporting can be&nbsp;&nbsp;

00:24:48.320 --> 00:24:54.640
really additive to risk management understanding.&nbsp;
And again, I think there's real opportunity here&nbsp;&nbsp;

00:24:54.640 --> 00:25:01.760
to look at this kind of reporting and understand,&nbsp;
you know, what's valuable for stakeholders and&nbsp;&nbsp;

00:25:01.760 --> 00:25:07.840
where is there opportunity to go further in really&nbsp;
informing value chains and policymakers and the&nbsp;&nbsp;

00:25:07.840 --> 00:25:14.960
public about AI risk and opportunity and what&nbsp;
can we learn again from other domains that have&nbsp;&nbsp;

00:25:14.960 --> 00:25:20.960
done this kind of work over decades to really&nbsp;
refine that kind of information sharing. 
 
 

00:25:20.960 --> 00:25:25.120
SULLIVAN: It's really great to hear about all the&nbsp;
advances that we're making on these reports. I'm&nbsp;&nbsp;

00:25:25.120 --> 00:25:30.640
guessing a lot of the metrics in there are&nbsp;
technical, but sociotechnical impacts—jobs,&nbsp;&nbsp;

00:25:30.640 --> 00:25:36.000
maybe misinformation, well-being—are&nbsp;
harder to score. What new measurement&nbsp;&nbsp;

00:25:36.000 --> 00:25:40.720
ideas are you excited about, and do you have any&nbsp;
thoughts on, like, who needs to pilot those? 
 
 

00:25:40.720 --> 00:25:46.560
CRAIG DECKARD: Yeah, it's an incredibly&nbsp;
interesting question that I think also&nbsp;&nbsp;

00:25:46.560 --> 00:25:50.640
just speaks to, you know, the breadth&nbsp;
of, sort of, testing and evaluation&nbsp;&nbsp;

00:25:50.640 --> 00:25:57.360
that's needed at different points along&nbsp;
that AI life cycle and really not getting&nbsp;&nbsp;

00:25:57.360 --> 00:26:01.840
lost in one particular kind of testing&nbsp;
or another pre- or post-deployment and&nbsp;&nbsp;

00:26:01.840 --> 00:26:05.520
thinking expansively about the risks that we're&nbsp;
trying to address through this testing. 
 
 

00:26:05.520 --> 00:26:09.600
You know, for example, even with the&nbsp;
UK's AI Security Institute that has&nbsp;&nbsp;

00:26:09.600 --> 00:26:12.880
just recently launched a new program, a new team,&nbsp;&nbsp;

00:26:12.880 --> 00:26:18.400
that's focused on societal resilience research.&nbsp;
I think it's going to be a really important area&nbsp;&nbsp;

00:26:18.400 --> 00:26:24.720
from a sociotechnical impact perspective to&nbsp;
bring some focus into as this technology is&nbsp;&nbsp;

00:26:24.720 --> 00:26:30.400
more widely deployed. Are we understanding&nbsp;
the impacts over time as different people&nbsp;&nbsp;

00:26:30.400 --> 00:26:35.040
and different cultures adopt and use this&nbsp;
technology for different purposes?  
 
 

00:26:35.040 --> 00:26:39.920
And I think that's an area where there really&nbsp;
is opportunity for greater public-private&nbsp;&nbsp;

00:26:39.920 --> 00:26:45.840
partnership in this research. Because we all&nbsp;
share this long-term interest in ensuring that&nbsp;&nbsp;

00:26:45.840 --> 00:26:53.200
this technology is really serving people and&nbsp;
we have to understand the impacts so that we&nbsp;&nbsp;

00:26:53.200 --> 00:26:58.880
understand, you know, what adjustments we can&nbsp;
actually pursue sooner upstream to address&nbsp;&nbsp;

00:26:58.880 --> 00:27:03.520
those impacts and make sure that this&nbsp;
technology is really going to work for&nbsp;&nbsp;

00:27:03.520 --> 00:27:07.480
all of us and in a way that is consistent&nbsp;
with the societal values that we want. 
 
 

00:27:07.480 --> 00:27:11.440
SULLIVAN: So, Amanda, looking ahead,&nbsp;
I would love to hear just what's going&nbsp;&nbsp;

00:27:11.440 --> 00:27:15.025
to be on your radar? What's top of&nbsp;
mind for you in the coming weeks? 
 
 

00:27:15.025 --> 00:27:19.920
CRAIG DECKARD: Well, we are certainly continuing&nbsp;
to process all the learnings that we've had from&nbsp;&nbsp;

00:27:19.920 --> 00:27:26.080
studying these domains. It’s really been a rich&nbsp;
set of insights that we want to make sure we,&nbsp;&nbsp;

00:27:26.080 --> 00:27:33.680
kind of, fully take advantage of. And, you&nbsp;
know, I think these hard questions and,&nbsp;&nbsp;

00:27:33.680 --> 00:27:40.480
you know, real opportunities to be thoughtful&nbsp;
in these early days of AI governance are not,&nbsp;&nbsp;

00:27:40.480 --> 00:27:45.920
sort of, going away or being easily resolved&nbsp;
soon. And so I think we continue to see value&nbsp;&nbsp;

00:27:45.920 --> 00:27:51.360
in really learning from others, thinking&nbsp;
about what's distinct in the AI context,&nbsp;&nbsp;

00:27:51.360 --> 00:27:56.000
but also what we can apply in terms of&nbsp;
what other domains have learned.  
 
 

00:27:56.000 --> 00:28:00.080
SULLIVAN: Well, Amanda, it has been such a&nbsp;
special experience for me to help illuminate&nbsp;&nbsp;

00:28:00.080 --> 00:28:05.120
the work of the Office of Responsible&nbsp;
AI and our team in Microsoft Research,&nbsp;&nbsp;

00:28:05.120 --> 00:28:08.000
and [MUSIC] it's just really special to&nbsp;
see all of the work that we're doing to&nbsp;&nbsp;

00:28:08.000 --> 00:28:12.000
help set the standard for responsible&nbsp;
development and deployment of AI. So&nbsp;&nbsp;

00:28:12.000 --> 00:28:18.560
thank you for joining us today, and thanks&nbsp;
for your reflections and discussion.  
 
 

00:28:18.560 --> 00:28:22.240
And to our listeners, thank you so much for&nbsp;
joining us for the series. We really hope you&nbsp;&nbsp;

00:28:22.240 --> 00:28:29.760
enjoyed it! To check out all of our episodes,&nbsp;
visit aka.ms/AITestingandEvaluation, and if you&nbsp;&nbsp;

00:28:29.760 --> 00:28:36.240
want to learn more about how Microsoft approaches&nbsp;
AI governance, you can visit microsoft.com/RAI. 
 
 

00:28:36.240 --> 00:28:49.840
See you next time! 
 
 

00:28:49.840 --> 00:28:50.670
[MUSIC FADES]

