TY - GEN
T1 - Quantifying the Privacy of Counterfactuals by Leveraging Membership Inference Attacks Against Synthetic Data
AU - Babaei, Maryam
AU - Wang, Yingke
AU - Lautraite, Hadrien
AU - Arcolezi, Héber H.
AU - Aïvodji, Ulrich
AU - Gambs, Sébastien
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/6/25
Y1 - 2026/6/25
N2 - Counterfactuals are typically used in high-stakes decision areas to explain a machine learning model by showing how changes to the user profiles result in the desired outcome. However, explaining the model's decisions through counterfactuals can also be exploited by an adversary to conduct privacy attacks against the model or its training data. Drawing on the analogy that counterfactuals provide realistic substitutes for real training data, similar to synthetic data, we demonstrate in this paper how it is possible to successfully perform privacy attacks on counterfactuals by drawing on the attacks developed against synthetic data. More precisely, we investigate the effectiveness of the membership inference attacks designed for synthetic data on various types of counterfactuals. Additionally, while existing membership inference attacks against counterfactuals usually require to be able to query the model, we show how it is possible to perform successful membership inference attacks using only a set of counterfactuals, with no access to the model from which they are generated. Our results demonstrate that model developers should be more cautious when releasing counterfactuals to various users, as it can lead to a privacy breach.
AB - Counterfactuals are typically used in high-stakes decision areas to explain a machine learning model by showing how changes to the user profiles result in the desired outcome. However, explaining the model's decisions through counterfactuals can also be exploited by an adversary to conduct privacy attacks against the model or its training data. Drawing on the analogy that counterfactuals provide realistic substitutes for real training data, similar to synthetic data, we demonstrate in this paper how it is possible to successfully perform privacy attacks on counterfactuals by drawing on the attacks developed against synthetic data. More precisely, we investigate the effectiveness of the membership inference attacks designed for synthetic data on various types of counterfactuals. Additionally, while existing membership inference attacks against counterfactuals usually require to be able to query the model, we show how it is possible to perform successful membership inference attacks using only a set of counterfactuals, with no access to the model from which they are generated. Our results demonstrate that model developers should be more cautious when releasing counterfactuals to various users, as it can lead to a privacy breach.
KW - Counterfactuals
KW - Membership inference attacks
KW - Privacy
KW - synthetic data
UR - https://www.scopus.com/pages/publications/105044415446
U2 - 10.1145/3805689.3812361
DO - 10.1145/3805689.3812361
M3 - Contribution to conference proceedings
AN - SCOPUS:105044415446
T3 - ACM FAccT 2026 - Proceedings of the 9th annual ACM Conference on Fairness, Accountability, and Transparency
SP - 4988
EP - 5022
BT - ACM FAccT 2026 - Proceedings of the 9th annual ACM Conference on Fairness, Accountability, and Transparency
PB - Association for Computing Machinery, Inc
T2 - 9th Annual ACM Conference on Fairness, Accountability, and Transparency, ACM FAccT 2026
Y2 - 25 June 2026 through 28 June 2026
ER -