Building AI, Responsibly

Equity-Centred and Contextual Design

Aubrey Blanche introduces equity-centred design and contextual fairness as foundations for responsible AI. Teams should assess impacts early, work with affected communities and define fairness for the system’s real domain.

Participation, Transparency and Accountability

Inclusive participation brings diverse perspectives into design, development, testing and governance. Plain-language explanations, accountability and audit trails make system decisions visible and reviewable.

Bias Mitigation and Human Safeguards

Blanche treats bias mitigation as continuous and measurable, supported by representative data and protected-attribute testing. High-stakes decisions require human review, escalation paths and trained reviewers.

Monitoring, Privacy and Data Dignity

Responsible AI continues after launch through monitoring, independent audits and lived-experience feedback. Privacy by design, meaningful consent and protection from exploitative data use preserve human dignity.

Hey, everyone. I'm Aubrey, white appearing woman with, pink, purple, and orange hair. Yes. Really, black turtleneck with gold accents. So I wanted to go in a different direction partially because I wouldn't stand up to comparison to Sarah, but I'm also gonna give you all of my IP in this talk. So, everybody have a camera on your phone?

Yeah? Yeah. This is the time to take photos. So I wanted to take us in a little bit of a different direction to talk about specifically what it means to to build AI responsibly. And if you know anything about my background, I will tell you that to me that means equitably. So being really thoughtful about the way that power dynamics operate in the world, but I wanna talk about the specific tactics that you can use to do that in line with my personal equitable AI framework that I use with my clients.

So first principle is equity centered design. So we actually must intentionally design AI systems to produce equitable outcomes. They will not do it on their own, it but turns out that your intentions actually are not good enough. And what I mean by that is just hoping that AI becomes more equitable is not a good way to get there.

So first, we need to conduct impact assessments at the point of conception of these tools. So we need to be thinking not only about what could go right in our ideal scenario, but what could go wrong, either accidentally or because nefarious actors get ahold of what we're doing. We need to be able to rank the potential outcomes, opportunities, and risks that we see by the likelihood and severity because the reality is you probably can't design against every bad that could ever happen.

But you probably can design against the most likely bads that would be really bad if they happened. We also wanna make sure that we're prioritizing user research with folks who are most likely to experience harm and fucking pay them for their time. Evaluate the outputs, not just for accuracy or efficiency or whatever the board of directors thinks AI can do this week, but actually look for equitable distribution of both benefits and harms.

We'll talk a little bit later about definitions of fairness, but it's really important to understand that the concept of fairness is very complicated. There is not actually one mathematical definition, and there's a lot of companies that have actually gotten into a lot of trouble because they've argued, hey. The error rate for people is the same, so it's fine. A really notable example of this, ProPublica in The US showed that there was a company called Northpoint.

Their particular algorithms were actually being used to predict recidivism amongst folks who had been convicted of crimes. Now the thing is their error rate for white and black people was the same. The problem is it was more likely to predict recidivism that didn't happen amongst black people and underestimated the likelihood of recidivism amongst white people.

So really fucked. Right? And so I think it's important that when you think about errors or you think about the way something impacts, we need to think about and really define that we're not having these divergent things. So next principle is contextual fairness. So fairness is not universal, and it doesn't have one meaning. So AI systems have to be tailored to the specific context that you're in, and you likely will have to make principled choices about what type of fairness you're optimizing for.

Because depending on the paper you're reading, there's somewhere between eight and thirty two mathematical definitions of fairness, and you cannot optimize for all of them at the same time. So you need to define the specific fairness metrics that are most impactful and important to your specific use case and validate them with domain experts and also ask affected users what they think.

Do you actually build credibility with the people whose lives are gonna be impacted by your technology? Now avoid just relying on the idea of statistical parity. Know that things can be wrong even if your large end math says it's fine. Right? Qualitative and quantitative experiences of fairness matter. And then document these choices.

So we know that model cards are an innovation that were actually created by AI ethicists to create more transparency about the way that models are trained and they perform. This is the type of information. Unfortunately, there are no standardizations about what have to go into model cards. If anyone like me has spent a lot of time with them for the foundational models, as you can imagine, Anthropic has pretty deep safety information In their model cards, Meta has less, as we would all imagine. But really thinking about being able to transparently define what the metrics that you're using are and be willing to share that with your users.

This can actually be a competitive advantage, especially if you're in the b to b space because I can tell you, still being an adviser to Culture Ramp, is that they often get questions about bias, but the askers don't actually know what answers they're looking for. So people know that AI is biased, but I would say that in the market, buyers are actually ill equipped to actually properly assess that. And so you can have a competitive advantage by having a strong point of view on what that means in your context.

Now inclusive participation. So across every aspect of design deployment and decommissioning, we should be having a diverse set of perspectives. We know this is a major problem. I've been working on the issue of diversity in tech since about 2013, and I will tell you it fucking sucks worse in AI. So the reason we need to think about that is because often when we're talking about just the engineering, the folks who are building AI, we're seeing worse skews in terms of demographics than we did in, say, classic software engineering.

But that doesn't mean diverse perspectives can't be represented in the process. So it turns out that AI ethicists, for example, are more female, more brown, more queer, and more disabled than AI researchers. And so making sure that ethicists are a part of the process, but also thinking of product managers, designers, anyone else involved in the process to try to look around and say, is a balanced team actually building this? Do we have the perspectives in the room, both defined by lived experience and professional expertise to provide well rounded feedback on the solutions that we're developing?

And it's also important that we make sure that those voices get into the room. So thinking about throughout your development process, what are your power balancing techniques that you have? So one, are you sending agendas for meetings in advance? Because neurodivergent people, it turns out neurotypicals also benefit from this, but it's especially helpful for introverts, both groups with executive functioning challenges.

Making sure that you're rotating decision authority or moving it around and giving people a chance to have input. Use RACIES so people know how their opinions are being considered. Offer anonymous feedback channels or when you're doing a brainstorm, make sure that you ask the question, give people time to collect their thoughts so the extrovert doesn't always win, speaking as the extrovert.

So making sure that the way that we're designing includes a variety of different people across the development process. Now transparency and accountability is critical. Now understanding that generative AI models are actually in many ways unexplainable. There's really interesting lines of research happening around that, but we also know that the models themselves don't always know how to report the way that reasoning has happened.

But to the best of our ability, we should be documenting our choices, and we should try to be accountable to a broad variety of stakeholders. So we need to publish plain language explanations of how these work. Again, model cards are a great way to do this. Even if there isn't a standardized format, this is something your company can innovate on to be able to look at how the systems work, the intended use, and also the limitations.

So that can include warnings about how not to use this, knowing that just putting it in a wildcard is not actually sufficient to get us to users fully understanding the system. So assign clear accountability for outcomes to humans. Be able to tell users who is accountable, whether it's them or whether it's someone at the company for those things.

Leveraging human in the loop, is incredibly important. So what we know is that AI is not actually smart enough to be making decisions of consequence about people. AI is not as smart as Sam Als Altman would have you believe. And I always say, remember, how many commas in his bank account depend on you believing the hype. Right? So maintain audit trails of training data, design decisions, and model updates for internal and external review.

Doing this, I always say, we can use AI to create documentation that does this doesn't need to be the most painful thing, But being able to document your decisions over time is helpful. Yes. If you're deploying in The US where everybody loves to hire a lawyer, that's important. But also because if something inevitably does go wrong with your system, we know that these these systems do fail, that you can look back and actually root cause why that happened so you can address it or prevent it in the future from happening again.

So as Sarah really beautifully talked about, AI reflects the social biases that are embedded in training data, but also understanding that biases can be introduced to your system by the end users. Right? So there are multiple sources of bias throughout the pipeline. I would be remiss if I did not mention that across the literature on AI, there is no shared and agreed upon definition of the word bias. So when you're talking about it, it's really important to be explicit about what you mean because the reality is not all bias is harmful.

So it for example, you could curate a a training set that is biased against the type of baseline problems that you see in training data. So maybe don't use Reddit as your training data because we know that there tends to be more male voices on Reddit, there also tends to be some nasty misogynistic shit on it. And so when we talk about bias, I find it very useful to specify whether we're talking about harmful bias or statistical skew or other definitions that we're using.

So being clear. But making sure that you're regularly testing your models for disparate impact across demographics or across populations of interest to your use case. Implementing bias mitigation strategies that prioritize not just equity of process, but ideally that you see equity of outcomes. So that doesn't mean forcing same outcomes where the data doesn't suggest that's actually what you should be doing, but thinking about as a proxy or a first assumption that equity of outcome tends to speak to a higher equity level in a process.

And also continuously update the models and retrain them with rep representative datasets as populations and context evolve. So I see this especially for companies that are moving into new markets globally, that often the models that they trained on one set of data are simply not appropriate for new cultural contexts. And so all of the things that I've talked about actually need to be redone because AI, for all of the things that is good at extrapolating out of sample, is a really risky thing to do. And so making sure that while we may have some learnings, say, from deploying in Australia or The US, if we're expanding across APAC, there are fundamentally different cultural norms and ways that data will be injected into your systems or the way that your existing models, what they're trained on is simply just not relevant to new markets.

So human oversight. So we talk about this AI, again, is not smart enough to be making decisions for you. So when you're thinking about proper uses of AI, we wanna think about judgment augmentation. So in this moment where content and reasoning is cheap, what is actually more valuable in this moment is human judgment and discernment. And so building in safety checks to AI that allows those humans to make contextual decisions.

So we said design human in the loop processes for critical decisions, but also being aware of the research on algorithmic acceptance. So what the research tells us is that the less expert someone So is in a given domain, the more likely that they are to accept the outcome of algorithms. So when you're designing human in the loop processes, it actually makes a huge difference when you specify the persona of who is acting as that human in the loop.

Right? So experts are more likely to reject the outcome of algorithms likely because they have access to a broader internal dataset. Right? There's also research that shows that those with higher levels of AI literacy and specifically in understanding how the technology itself works are less likely to blindly trust AI. So people are also less likely to use AI the more that they know about it, which I always find a fascinating statistic.

But so thinking about any human is generally not good enough. There is some very nascent research showing that middle aged women are actually the best at making judgement AI augmented decisions. And the reason for that is because they're actually more likely to challenge algorithmic outcomes than any other demographic group. So I like that.

Like, right, you get to middle age, you just take no shit even from the technology. Makes me really excited to be close to my forties. Human oversight. Making sure that you're creating escalation protocols for edge cases. So this means that your users, when they log in to your system, should know how to call you if something is going wrong.

So if you say, oh, they'll just put it in the support line. I'm telling you that is not good enough. Someone actually has to know where they can go, and you need to have an escalation process for issues of algorithmic harm. I always tell my clients to operate from the point of view of probability of harm equals one.

Your ability to anticipate it is not one, And so you need to be aware that something may happen outside your control or anticipation, and what you wanna do is know about that as quickly as possible so that you can rectify the situation. Also, provide training for human reviewers. So going back to that idea of the more expert someone is, the more qualified and capable they are to actually act as an effective human in the loop.

So all of this has to be done all of the time. And I know that's probably not what your CFO wants to hear, but it is true. So what we know is that this technology evolves and changes over time, and so there is no rotisserie chicken, set it and forget it version of developing AI. So established metrics, ideally dashboards, if you have the instrumentation and observability practices to monitor system performance across things like equity, fairness, safety in the same context that you model other types of algorithmic performance.

So you should not have an equity dashboard and an everything else dashboard. You should be looking at the overall health of the system and considering equity and fairness in that. Schedule recurring bias audits, or if you're fancy, do continuous monitoring on at least a sample of outputs to understand if model drift is happening from the original specifications and testing that you did.

And also build user feedback loops. So regular conversations with users, again, especially those who are most likely to be experiencing harm, to understand the lived experience. This is also just gonna make you a better Right? Like, not a lot of this sounds like it's something special we do for social justice. It actually just improves the outcomes that these products are creating.

So last, of course, we wanna think about safety and privacy, but I really wanna bring in this idea of data dignity into this. So, certainly, you know, your CISA wants you to care about this, but I think it's really important why do we care about safety. Because someone's individual identity and their self is incredibly important and worthy of respect, And one of the ways that we hold that sacred is through respecting general principles and regulations around safety and privacy.

So applying privacy by design principles, things that you're probably already subject to according to your general counsel, things like minimization, anonymization, and making sure that storage data and transit is secure. Get explicit consent for everything. And if you're writing that, like, 4,000 word t and c document, I'm telling you, that's not actually informed consent even if your lawyer would tell you it's sufficient.

Right? Law and ethics are actually not the same thing. And prohibit data uses that may stigmatize, exploit, or endanger vulnerable groups and be willing to update your practices if you get feedback that you're doing that unintentionally. Right? So it can happen. A lot of us in this room come from less marginalized or intersectionally privileged backgrounds, and so that mean we may simply lack the lived experience to anticipate all of the ways something go wrong.

It's okay if something goes wrong, but make sure that you fix it. So that was all I have for you today. If you want the slides, you can get me on LinkedIn. But, hopefully, this has given you some meaningful and really specific actions that you can take. What I would say is I've given you a lot. That was 24 different actions that you could take, but what I would say, the most useful question that I hope you get out of this is what's one or two of these things that I know I can get into our process in the next week and start there. Start building those practices over time because they feed on each other and actually become easier.

So thank you so much for your time. It was great to be here.

Building AI, Responsibly

The Mathpath

Equity-Centered Design

AI systems must be designed to actively promote equitable outcomes, prioritising the needs and experiences of historically marginalised communities.

  • Conduct equity impact assessments at the start of every AI project
  • Prioritise user research with communities most at risk of harm or exclusion
  • Evaluate outputs for equitable distribution of benefits and burdens

Contextual Fairness

Fairness is not universal; AI systems must be tailored to context to avoid reinforcing systemic inequities.

  • Define fairness metrics specific to the system’s domain and validate them with domain experts and affected users
  • Avoid relying solely on statistical parity
  • Document contextual fairness choices

Inclusive Participation

AI development must include diverse perspectives across design, development, testing and governance.

  • Involve community representatives and advocacy groups in design reviews
  • Include members from diverse demographic, cultural and disciplinary backgrounds
  • Use structured power-balancing practices so all voices are heard

Transparency & Accountability

AI systems must be explainable, documented and accountable to users, stakeholders and regulators.

  • Publish plain-language explanations of how systems work and their limitations
  • Assign clear accountability for AI outcomes
  • Maintain audit trails of training data, design decisions and model updates

Bias Recognition & Active Mitigation

AI systems inherently reflect societal biases; mitigation must be intentional, continuous and measurable.

  • Test regularly for disparate impacts across protected attributes
  • Implement bias-mitigation strategies prioritising equity of outcomes
  • Update models with more representative datasets as contexts evolve

Human Oversight & Safeguards

AI should augment, not replace, human judgment—especially in high-stakes contexts.

  • Design human-in-the-loop review for critical decisions
  • Create escalation protocols for edge cases or uncertain outputs
  • Train human reviewers to identify and correct potential harms

Continuous Monitoring & Evaluation

Responsible AI is an ongoing practice requiring regular evaluation, iteration and correction.

  • Monitor performance across equity, fairness and safety dimensions
  • Schedule recurring bias audits, including external independent reviews
  • Build feedback loops to capture lived experiences and incorporate findings

Safety, Privacy & Data Dignity

Data use in AI must respect individual dignity, safeguard privacy and prevent harm.

  • Apply privacy-by-design principles
  • Seek explicit consent for sensitive data use
  • Prohibit data uses that may stigmatise, exploit or endanger vulnerable groups

Thank you!

A portrait of Aubrey Blanche, wearing a bright yellow jacket, appears within the colourful geometric design.

Concepts & Methods

  • equity-centred design
  • contextual fairness
  • inclusive participation
  • AI accountability
  • bias mitigation
  • human oversight
  • continuous monitoring
  • privacy by design
  • data dignity