Cover image for “The Data Dilemma: Reconciling Privacy and Fairness in AI”
← essays · 2025-02-10

The Data Dilemma: Reconciling Privacy and Fairness in AI

Privacy and fairness are presented as opposing values in AI governance, but the choice between them is not actually a binary. A close look at GDPR exceptions, US self-regulation, the Amazon hiring algorithm, the Apple Card, and IBM's facial-recognition exit.

Originally written for AI Ethics and Society · University of Cambridge · February 2025 · ~4,400 words

I. Introduction

Artificial intelligence (AI) systems are expanding across diverse domains, from credit scoring and hiring to public services, raising important questions about which values AI governance should prioritise. On one side is privacy, long recognised as fundamental for individual autonomy and dignity. On the other side is fairness, referring to the imperative that AI systems do not replicate or amplify unjust discrimination (Calo, 2017). As AI increasingly determines access to opportunities and resources, the challenge of striking an equitable balance between privacy and fairness has never been more urgent.

These two values are sometimes presented as starkly opposed. Organisations mindful of privacy may be reluctant to collect sensitive data such as ethnicity or sexual orientation, fearing both ethical concerns and regulatory penalties. Yet without such information, it can be harder to detect whether an AI system is systematically disadvantaging a protected group, a tension increasingly noted by policymakers (Gupta, 2023). This conflict is powerfully illustrated by van Bekkum & Borgesius (2022), who argue that under the General Data Protection Regulation (GDPR), the “special category” rules restrict processing data on, for example, ethnicity. While well-intentioned to protect citizens’ privacy, these restrictions can inadvertently hamper an organisation’s ability to audit and correct AI-driven discrimination. Conversely, if the pendulum swings too far towards fairness objectives, a system might collect and process more intrusive personal data to root out hidden biases, risking large-scale surveillance or data misuse.

The consequences of such prioritisation choices are far-reaching. If privacy takes precedence to the exclusion of collecting sensitive attributes, communities can suffer silent injustices: biased algorithms go unrecognised because there are no reliable means to measure disparate impact. Yet if fairness is pursued without respect for privacy, individuals may be subjected to invasive data gathering that endangers their security and autonomy. While these tensions appear entrenched, they need not be irreconcilable. Careful governance (through legislation, corporate policy, and technical safeguards) can uphold both values, provided stakeholders acknowledge the trade-offs and adopt context-appropriate approaches.

The rest of this essay explores how prioritising privacy or fairness shapes AI’s societal impact and ethical landscape. Section II defines privacy and fairness in AI, examining why the two often clash. Section III surveys the EU’s relatively strict data protection regime against the US’s market-driven approach, showing how each system embeds or compromises these values. Section IV examines three case studies to illustrate the consequences of governance choices, weaving in the alternative governance frameworks (relational data governance, data trusts) that emerge as the most promising response. The Conclusion reiterates the risks of an unbalanced approach.

II. Defining the Tension: Privacy vs. Fairness in AI

A core challenge in AI governance is reconciling individuals’ rights to privacy with the collective imperative for fairness. These values, while both fundamental, are often positioned in tension. This tension becomes particularly acute when AI systems, designed to optimise outcomes, require personal data to function effectively and be audited for discriminatory impacts.

Privacy: From Individual Right to Data Protection Principles

Privacy has traditionally been understood as an individual’s “right to be let alone” (Warren & Brandeis, 1890). While individual informational self-determination (the ability to control how one’s data is collected and used; Westin, 1967) remains a core tenet, AI’s reliance on vast datasets and complex algorithms introduces collective privacy dimensions. The aggregation and analysis of data, even when seemingly innocuous at the individual level, can reveal sensitive patterns about groups and communities (Solove, 2008; Dencik et al., 2019).

Over time, especially within European legal frameworks, this concept has evolved from an individual right into a comprehensive data protection system as well. The GDPR embodies this evolution, enshrining principles such as data minimisation (Article 5), conditions for explicit consent (Article 6), and restrictions on processing sensitive data (Article 9), including information on ethnicity, religion, and sexual orientation. Butterworth (2018) highlights that while the GDPR attempts to balance these principles with fairness requirements, its structure still prioritises privacy over the proactive use of sensitive attributes for bias detection.

Fairness: Formal Equality vs. Substantive Equity

Fairness in AI centres on preventing or remediating group-based discrimination, encompassing a spectrum of interpretations. Formal equality dictates that AI systems must not explicitly use protected attributes like race or gender in decision-making (Barocas & Selbst, 2016). This “attribute-blind” approach aims to prevent direct discrimination. However, Barocas and Selbst (2016) critique that “fairness through unawareness” fails to address indirect discrimination, where seemingly neutral features act as proxies for protected attributes or where structural inequalities are embedded in the training data.

An alternative perspective, substantive equality, focuses on equitable outcomes, regardless of whether protected attributes are explicitly used. This aligns with anti-disparate impact doctrine, where a system is deemed discriminatory if it produces systematically worse outcomes for protected groups, even without overt bias in its design. In practice, AI systems are susceptible to emergent biases, where seemingly benign features like postal codes correlate strongly with protected attributes like race and socio-economic status, or where historical disparities are embedded into training data (Barocas and Selbst, 2016; Hoffmann, 2019). Consequently, standard “colour-blind” design can still perpetuate inequality if underlying patterns remain unexamined.

The Inherent Clash: Competing Imperatives

The tension between privacy and fairness in AI arises from fundamentally conflicting imperatives:

  1. Data Minimization vs. Data Access: The GDPR’s principle of data minimisation clashes directly with the data needs of fairness auditing. Detecting disparate impacts often requires access to sensitive attributes to identify and measure outcome disparities across groups. Van Bekkum and Borgesius (2022) and Gupta et al. (2023) highlight that restrictions on processing special categories of data leave developers without the necessary demographic data to conduct thorough bias assessments. Although the European Union’s AI Act’s Article 10(5) allows for narrowly defined exceptions where providers of high-risk AI systems may process special categories of personal data to detect and correct biases, these exceptions are not all-encompassing and leave developers uncertain about the legality of collecting protected attributes while operating outside the EU.
  2. Individual Control vs. Collective Harms: The GDPR’s emphasis on individual consent mechanisms is insufficient to address group-level harms. Niklas and Dencik (2024) argue that individual choices, such as opting out of data collection, cannot prevent systemic discrimination that affects entire demographics.
  3. The Proxy Problem: Even without the explicit use of protected attributes, “proxy discrimination” can occur when seemingly neutral features correlate strongly with protected characteristics (Barocas & Selbst, 2016). Removing these proxies may reduce accuracy and may not resolve the issue, shifting the burden of the problem rather than dissolving it.
  4. Transparency Paradox: While transparency is vital for identifying bias, it also risks exposing personal data or proprietary model weights. Organisations concerned with protecting their intellectual property, or obfuscating their use of others’ intellectual property, may resist full disclosure, hindering external scrutiny (Butterworth, 2018).

Societal Stakes

Where privacy requirements become overly restrictive, data controllers may lack the information needed to spot biased patterns, enabling discrimination to continue undetected. Meanwhile, an uncritical pursuit of fairness metrics, unconstrained by privacy or data protection, can legitimise intrusive data collection, ironically subverting individuals’ rights. Both extremes can erode public trust as citizens are alienated by excessive surveillance or frustrated by covert bias.

III. Governance Mechanisms and Their Consequences

Debates about privacy and fairness in AI governance are frequently embodied in contrasting approaches: the European Union’s stringent data protection regime and the United States’ market-driven, self-regulatory model. Each presents distinct trade-offs in addressing algorithmic discrimination while protecting individual rights.

The EU Approach: Privacy as Paramount

The EU’s GDPR prioritises privacy through data minimisation and strict limitations on processing “special categories” of data: race, ethnicity, health information, religion. While designed to limit invasive data practices, this framework can inadvertently hinder efforts to detect and mitigate systemic biases. Collecting and processing protected characteristics is generally prohibited unless explicit exceptions apply.

The EU AI Act attempts to address this through a risk-tiered approach: “high-risk” AI systems face stricter obligations on data quality, transparency, and human oversight. Article 10(5) permits processing special-category data when “strictly necessary” for bias detection. Critics argue this exception is too narrow and, combined with the GDPR’s data minimisation principle, may still deter thorough audits (van Bekkum & Borgesius, 2022). The Act’s framing of discrimination as a “risk” to be managed rather than a structural issue embedded in social and economic power dynamics has drawn additional criticism: Niklas and Dencik (2024) argue this risk-based framing fails to address the underlying market logics that produce discriminatory outcomes.

By treating privacy as non-negotiable, the EU ensures citizens’ personal data is not gathered indiscriminately. This builds public trust and encourages acceptance of AI under strict rules. Nonetheless, structural discrimination can remain hidden. Smaller organisations hesitate to collect sensitive data, leaving them ill-equipped to identify biased outcomes. Vulnerable groups may thus lack clear evidence of disparate impact.

The US Approach: Innovation and Fairness as Market-Led

In contrast to the EU, the United States lacks a comprehensive federal privacy law comparable to the GDPR, resulting in a more permissive environment for data collection (Calo, 2017). Outside of a patchwork of industry- and state-level laws, the absence of stringent data protection regulations theoretically facilitates robust bias detection through access to larger, more diverse datasets. However, this data abundance carries significant risks: reliance on corporate self-regulation can lead to excessive surveillance and unchecked discriminatory practices (Calo, 2017). Individuals often lack meaningful legal recourse if they suspect bias or data misuse.

Under the second Trump administration, AI policy has shifted toward reducing regulatory barriers. An executive order signed on January 23, 2025, aims to eliminate previous policies alleged to be hindrances to AI development, emphasising the need to sustain and enhance America’s global AI dominance (The White House, 2025). This deregulatory approach may further exacerbate concerns regarding data privacy and the potential for discriminatory practices. In the absence of federal oversight, individual states have begun to enact their own AI and privacy regulations, leading to a fragmented governance structure that creates inconsistencies for organisations operating across state lines (Fazlioglu, 2024).

In the absence of comprehensive federal rules, private ethics boards and voluntary guidelines have emerged. Major tech firms like Google, Microsoft, and IBM publish “Responsible AI” frameworks, while non-profit trade associations issue guidance on responsible AI adoption. These efforts can cultivate transparency and encourage baseline standards. However, critics warn of “ethics washing” (Hoffmann, 2019): companies may adopt feel-good principles without consistent, enforceable follow-through. Smaller entities may ignore such frameworks entirely. Without robust external enforcement, corporate boards or advisory councils can be overruled when profit motives predominate.

Comparing the EU and US Models

Both frameworks grapple with the same fundamental tension. The EU’s emphasis on data minimisation can hinder bias detection. The US’s reliance on corporate self-regulation risks unchecked data collection and inconsistent fairness efforts. Neither system fully resolves the dilemma. Reconciling fairness with privacy requires a balanced governance architecture that upholds data protection while enabling monitored audits of sensitive attributes. As the case studies below will show, this demands rethinking the consent model itself.

IV. Corporate Governance: Three Cases, and Why They Point Toward Relational Data Governance

The tension between privacy and fairness manifests in real-world deployments. The three cases that follow (Amazon’s hiring tool, the Apple Card approval algorithm, and IBM’s facial recognition technology) demonstrate how prioritising or neglecting certain values shapes ethical outcomes and public trust. Each case illustrates how different governance contexts produce different but related failures. Together they make the case for a different way of thinking about data governance entirely.

Amazon’s AI Hiring Tool: Data Abundance, But Not Fairness

Around 2014, Amazon developed an AI-driven recruitment tool to automate the screening of job applications for technical roles. The system, trained on a decade of historical hiring data, quickly learned to discriminate against female candidates. The training data reflected a male-dominated tech workforce, and the algorithm, without explicit gender labels, identified and amplified gendered language as proxies for quality (Simonite, 2018). The AI learned that male-centric patterns were successful, thereby reproducing past bias.

This case highlights how biased historical data perpetuates inequities. Amazon’s tool, trained on resumes reflecting the male-dominated tech industry, did not explicitly use gender as a variable but relied on proxies like language patterns and educational institutions, penalising female applicants (Dastin, 2018). While anti-discrimination laws like Title VII of the Civil Rights Act prohibit employers from using gender in hiring decisions, this legal safeguard may have inadvertently limited Amazon’s ability to test the model for disparate impacts across gender. If gender data had been appended to testing datasets, these biases might have been revealed and mitigated earlier.

This friction between privacy and fairness underscores a critical challenge: protecting sensitive attributes is essential for individual rights, yet their absence in testing can obscure systemic discrimination. Amazon’s failure to conduct thorough fairness audits allowed these biases to persist unchecked. Scrapping the tool demonstrates the risks of relying solely on internal accountability mechanisms without external oversight.

Apple Card Algorithm: Privacy as a Shield for Bias?

In 2019, the Apple Card, issued by Goldman Sachs, faced accusations of gender discrimination. High-profile users, including tech entrepreneur David Heinemeier Hansson, reported that their wives received significantly lower credit limits despite shared finances and comparable or superior credit scores (Vigdor, 2019).

Apple and Goldman Sachs emphasised that gender was not explicitly collected or considered by the algorithm, citing privacy and regulatory compliance in the U.S. financial services industry (Knight, 2019). This privacy-minded approach hindered the ability to detect algorithmic bias. Because the algorithm did not track gender, it was not straightforward to test systematically whether women received lower offers. Moreover, Goldman Sachs invoked proprietary secrecy, limiting external scrutiny of the model’s weighting of variables. If the model contained latent proxies for gender, this could result in a system that disadvantaged women (New York State Department of Financial Services, 2021).

This case highlights the paradox that prioritising privacy (or at least the appearance of privacy through data minimisation) can mask underlying bias. The lack of transparency, justified by privacy concerns and proprietary interests, eroded trust and prompted a regulatory investigation. Formal regulatory oversight came reactively, post-complaint, rather than preventing the incident.

IBM’s Facial Recognition: Beyond Privacy and Fairness

IBM’s experience with facial recognition provides a contrasting example. Initially, IBM developed and marketed facial recognition for various applications, including law enforcement. The 2018 “Gender Shades” study revealed significant accuracy disparities for darker-skinned women (Buolamwini & Gebru, 2018). This exposed how biased training data, often skewed toward lighter-skinned faces, leads to discriminatory outcomes.

IBM attempted to mitigate the bias by releasing a more diverse “Diversity in Faces” dataset (IBM Research, 2019). However, concerns persisted regarding potential misuse for mass surveillance, racial profiling, and human rights violations. In 2020, IBM took the significant step of exiting the general-purpose facial recognition market, stating it opposed “the uses of any technology […] for mass surveillance, racial profiling, [or] violating human rights” (Krishna, 2020). This decision reflects a recognition that some technologies, even if “fairer,” may be inherently incompatible with fundamental values.

What These Cases Point Toward: Relational Data Governance

Amazon’s case shows that data abundance without comprehensive fairness oversight amplifies existing biases. The Apple Card example highlights how gender-blindness combined with proprietary opacity can mask and entrench discrimination. IBM’s exit demonstrates that even technical fairness improvements may be insufficient when the underlying technology raises deeper ethical concerns.

These three failures share a common structure: each treats data as an individual property right or proprietary asset, and each treats fairness auditing as an internal corporate function. The result is that the people whose data is being used (and whose outcomes are being shaped) have no voice in either the governance of their data or the design of the fairness tests applied to it.

This is precisely the diagnosis that motivates Salomé Viljoen’s (2021) relational data governance framework. Viljoen argues that the standard liberal model, where data is personal property governed by individual consent, fundamentally misunderstands what data is. Even processing an individual’s data reveals patterns and risks that extend to peers, families, and entire communities. A purely individualistic consent model cannot capture these broader societal consequences. Similarly, a rigid privacy-centric approach, while protective of individual autonomy, may hinder legitimate fairness audits or socially valuable research.

Relational governance reconciles these competing interests by foregrounding collective decision-making about data use. Data shapes individual outcomes, group-level dynamics, and power relations simultaneously. Under a relational model, data subjects would collectively authorise specific forms of sensitive data collection, including protected attributes for bias auditing. This contrasts sharply with the prevailing emphasis on individual consent, which, as Niklas and Dencik (2024) highlight, often obscures systemic discrimination.

Data trusts offer a concrete instantiation. These structures establish a fiduciary relationship where data is managed on behalf of a community or group, with clearly defined purposes and oversight. The UK Biobank serves as a data trust by collecting and managing health data from approximately 500,000 participants who consent to their data being used for medical research. This model allows the collection of sensitive demographic data under strict conditions and community control, mitigating privacy risks while facilitating bias detection. By placing data stewardship in trusted entities accountable to the communities they serve, data trusts can uphold collective interests, mitigate privacy risks, and facilitate fairness audits.

Privacy-Enhancing Technologies (PETs) complement this institutionally. Differential privacy, for instance, introduces statistical noise to datasets, allowing meaningful analysis without compromising individual identities. Integrating PETs into data governance models strengthens their ability to balance data utility with robust privacy safeguards.

Relational governance is not a complete answer. Establishing legitimate representation and decision-making processes within diverse communities is complex. Scale, stakeholder engagement, and accountability remain significant hurdles. Determining which groups qualify as relevant data stakeholders and resolving internal disagreements fairly poses ongoing challenges.

But applied to the three cases above, the framework reveals what is missing. Amazon’s hiring tool was governed entirely by Amazon; no collective body of job applicants, technologists, or labour advocates had standing to audit it. The Apple Card algorithm was governed entirely by Goldman Sachs; no body representing cardholders had access to the model. IBM’s facial recognition was governed entirely by IBM until the company chose to exit; the affected communities (over-policed populations subject to facial recognition surveillance) had no governance role at all. Relational governance moves the question from “did the company comply with privacy law” to “who has standing to decide how this data is used, and what audits they can demand.”

V. Conclusion: Reflections on Prioritising Values in AI Governance

This essay has argued that a fundamental tension exists between privacy and fairness in the governance of AI that existing regulatory frameworks and corporate policies struggle to reconcile. The pursuit of one value, when prioritised to the exclusion of the other, can inadvertently compromise both. The consequences of these choices manifest in how AI systems distribute benefits and harms, often amplifying systemic inequalities when governance mechanisms fail to strike a balance.

The conceptual conflict between privacy (centred on individual autonomy and data minimisation) and fairness (focused on equitable outcomes and bias mitigation) is embedded within different governance approaches. The EU’s GDPR, emphasising data minimisation and restrictions on processing sensitive data, can hinder the data collection necessary for informative bias audits. The US’s more permissive, market-driven approach facilitates bias detection but also risks unchecked surveillance and harmful misuse of personal data.

The case studies demonstrate the real-world consequences of these trade-offs. Amazon’s experience shows how non-collection of data, even when driven by civil rights protections, can allow historical biases to persist unnoticed. The Apple Card algorithm demonstrates how an emphasis on privacy without adequate transparency can obscure discrimination. IBM’s withdrawal from facial recognition highlights that even AI accuracy improvements can raise broader concerns. These cases reveal the limitations of voluntary corporate policies and the challenges of addressing fairness within privacy-focused regulatory frameworks.

The alternative frameworks integrated into Section IV, relational data governance and data trusts, offer a conceptual shift. By emphasising the inherently collective implications of data processing and moving beyond individualistic notions of privacy and control, they advocate for shared decision-making and collective authorisation that current EU and US frameworks both lack.

Rather than offering a definitive solution, these findings underscore the need for proactive enforcement and auditing mechanisms that do not rely solely on either rigid data restrictions or industry self-regulation. The current governance landscape often fails to intervene until harm has already occurred. Alternative governance models, combined with technical tools like PETs, offer potential pathways for mitigating these tensions, but they require enforceable oversight rather than voluntary compliance.

An unbalanced approach to AI governance, prioritizing either privacy or fairness to the exclusion of the other, carries significant risks. While privacy and fairness may appear inherently at odds in the context of AI, the pursuit of thoughtful, adaptable, and proactively enforced governance offers the possibility of adapting both fundamental values to these new challenges.

References

  • Barocas, S., & Selbst, A. D. (2016). Big data’s disparate impact. California Law Review, 104, 671–732.
  • Buolamwini, J., & Gebru, T. (2018). Gender Shades. PMLR, 81.
  • Butterworth, M. (2018). The ICO and AI. Computer Law & Security Review, 34(2).
  • Calo, R. (2017). Artificial Intelligence Policy. UC Davis Law Review, 51, 399–435.
  • Dastin, J. (2018, October 10). Amazon scraps secret AI recruiting tool that showed bias against women. Reuters.
  • Dencik, L., Hintz, A., & Redden, J. (2019). Exploring data justice. Information, Communication & Society, 22(7).
  • Fazlioglu, M. (2024). US State Comprehensive Privacy Laws Report. IAPP.
  • General Data Protection Regulation. (2016). Regulation (EU) 2016/679.
  • Gupta, A., Ho, D. E., King, J., Webley-Brown, H., & Wu, V. Y. (2023). The Privacy-Bias Tradeoff. FAccT ‘23, 492-505.
  • Hoffmann, A. L. (2019). Where fairness fails. Information, Communication & Society, 22(7).
  • IBM Research. (2019). Diversity in Faces.
  • Knight, W. (2019, November 19). The Apple Card didn’t ‘see’ gender. WIRED.
  • Krishna, A. (2020, June 8). IBM CEO’s letter to Congress on racial justice reform.
  • New York State Department of Financial Services. (2021). Report on Apple Card investigation.
  • Niklas, J., & Dencik, L. (2024). Data justice in the “twin objective” of market and risk. Policy & Internet, 16(2).
  • Simonite, T. (2018). Amazon ditched AI recruitment software. MIT Technology Review.
  • Solove, D. J. (2008). Understanding Privacy. Harvard University Press.
  • UK Biobank. (n.d.). About UK Biobank.
  • van Bekkum, M., & Borgesius, F. Z. (2022). Using sensitive data to prevent discrimination by AI. Computer Law & Security Review, 48, 105770.
  • Vigdor, N. (2019). Apple’s credit card is under investigation. NYT.
  • Viljoen, S. (2021). A Relational Theory of Data Governance. Yale Law Journal, 131(2), 573–654.
  • Warren, S. D., & Brandeis, L. D. (1890). The right to privacy. Harvard Law Review, 4(5).
  • Westin, A. F. (1967). Privacy and Freedom. Atheneum.
  • The White House. (2025, January 23). Fact sheet: President Donald J. Trump takes action to enhance America’s AI leadership.