AI & ML interests
AI Made in Switzerland – Shaped by You Join the movement! Swiss {ai} Weeks calls on researchers, developers, businesses, and citizens to come together and build the future of AI — hands-on, ethical, and open. This isn't just exploration, it's collaboration in action.
Recent Activity
View all activity
jstuker
updated a
Space 10 months ago
hyunjun1121
authored a
paper 11 months ago
kechrisc
authored 4
papers 12 months ago
KID-PPG: Knowledge Informed Deep Learning for Extracting Heart Rate from a Smartwatch
Paper • 2405.09559 • Published
Don't Think It Twice: Exploit Shift Invariance for Efficient Online Streaming Inference of CNNs
Paper • 2408.03223 • Published
Time series saliency maps: explaining models across multiple domains
Paper • 2505.13100 • Published
DC is all you need: describing ReLU from a signal processing standpoint
Paper • 2407.16556 • Published
JackDapid
updated a
Space 12 months ago
JackDapid
published a
Space 12 months ago
Post
2041
🛡️ At Ai4Privacy, our goal is to empower researchers to build a safer AI ecosystem. Today, we're highlighting crucial research that does just that by exposing a new vulnerability.
The paper "Forget to Flourish" details a new model poisoning technique. It's a reminder that as we fine-tune LLMs, our anonymization and privacy strategies must evolve to counter increasingly sophisticated threats.
We're proud that the Ai4Privacy dataset was instrumental in this study. It served two key purposes:
Provided a Realistic Testbed: It gave the researchers access to a diverse set of synthetic and realistic PII samples in a safe, controlled environment.
Enabled Impactful Benchmarking: It allowed them to measure the actual effectiveness of their data extraction attack, proving it could compromise specific, high-value information.
This work reinforces our belief that progress in AI security is a community effort. By providing robust tools for benchmarking, we can collectively identify weaknesses and build stronger, more resilient systems. A huge congratulations to the authors on this important contribution.
🔗 Read the full paper: https://arxiv.org/html/2408.17354v1
#OpenSource #DataPrivacy #LLM #Anonymization #AIsecurity #HuggingFace #Ai4Privacy #World's largest open privacy masking dataset
The paper "Forget to Flourish" details a new model poisoning technique. It's a reminder that as we fine-tune LLMs, our anonymization and privacy strategies must evolve to counter increasingly sophisticated threats.
We're proud that the Ai4Privacy dataset was instrumental in this study. It served two key purposes:
Provided a Realistic Testbed: It gave the researchers access to a diverse set of synthetic and realistic PII samples in a safe, controlled environment.
Enabled Impactful Benchmarking: It allowed them to measure the actual effectiveness of their data extraction attack, proving it could compromise specific, high-value information.
This work reinforces our belief that progress in AI security is a community effort. By providing robust tools for benchmarking, we can collectively identify weaknesses and build stronger, more resilient systems. A huge congratulations to the authors on this important contribution.
🔗 Read the full paper: https://arxiv.org/html/2408.17354v1
#OpenSource #DataPrivacy #LLM #Anonymization #AIsecurity #HuggingFace #Ai4Privacy #World's largest open privacy masking dataset
Post
1158
In data privacy, 92% accuracy is not an A-grade. Privacy AI needs to be better.
That's the stark takeaway from a recent benchmark by Diego Mouriño
(Making Science), who put today's top PII detection methods to the test on call center transcripts using the Ai4Privacy dataset.
They pitted cutting-edge LLMs (like GPT-4 & Gemini) against traditional systems (like Cloud DLPs). The results show that our trust in these tools might be misplaced.
📊 The Hard Numbers:
Even top-tier LLMs peaked at a reported 92% accuracy, leaving a potential dangerous 8% gap where your customer's data can leak. They particularly struggled with basics like 'last names' and 'street addresses'.
The old guard? Traditional rule-based systems reportedly achieved a shocking 50% accuracy. A coin toss with your customers' privacy.
This tells us that for privacy tasks, off-the-shelf accuracy is a vanity metric. The real metric is the cost of a single failure—one leaked name, one exposed address.
While no tool is perfect, some are better than others. Diego’s full analysis breaks down which models offer the best cost-to-accuracy balance in this flawed landscape. It's a must-read for anyone serious about building trustworthy AI.
#DataPrivacy #AI #LLM #RiskManagement #MetricsThatMatter #InfoSec
Find the full post here:
https://www.makingscience.com/blog/protecting-customer-privacy-how-to-remove-pii-from-call-center-transcripts/
Dataset:
ai4privacy/pii-masking-400k
That's the stark takeaway from a recent benchmark by Diego Mouriño
(Making Science), who put today's top PII detection methods to the test on call center transcripts using the Ai4Privacy dataset.
They pitted cutting-edge LLMs (like GPT-4 & Gemini) against traditional systems (like Cloud DLPs). The results show that our trust in these tools might be misplaced.
📊 The Hard Numbers:
Even top-tier LLMs peaked at a reported 92% accuracy, leaving a potential dangerous 8% gap where your customer's data can leak. They particularly struggled with basics like 'last names' and 'street addresses'.
The old guard? Traditional rule-based systems reportedly achieved a shocking 50% accuracy. A coin toss with your customers' privacy.
This tells us that for privacy tasks, off-the-shelf accuracy is a vanity metric. The real metric is the cost of a single failure—one leaked name, one exposed address.
While no tool is perfect, some are better than others. Diego’s full analysis breaks down which models offer the best cost-to-accuracy balance in this flawed landscape. It's a must-read for anyone serious about building trustworthy AI.
#DataPrivacy #AI #LLM #RiskManagement #MetricsThatMatter #InfoSec
Find the full post here:
https://www.makingscience.com/blog/protecting-customer-privacy-how-to-remove-pii-from-call-center-transcripts/
Dataset:
ai4privacy/pii-masking-400k
Post
2731
Started
aistatuscodes as a new project to create codes to understand AI performance better.
Going to be posting daily here and on instagram until we get to 100m downloads :)
https://www.instagram.com/MikeDoesDo/
Follow along the journey!
Going to be posting daily here and on instagram until we get to 100m downloads :)
https://www.instagram.com/MikeDoesDo/
Follow along the journey!
Post
1565
PII-Masking-1M Final Day (7/7)! 🚀 Today, we unveil 5 NEW Enterprise PII (E-PII) Dataset PREVIEWS!
Standard PII tools often miss sensitive *business* data. That's why we built E-PII previews for the data that powers your operations and compliance needs.
Get a first look (representing 100,000 samples each!) into datasets designed for real-world enterprise security across these categories:
🏥 **PHI Preview**: For Healthcare Data
💳 **PFI Preview:** For Financial Data
🏢 **PWI Preview:** For Workplace Data
💻 **PDI Preview:** For Digital Activity Data
📍 **PLI Preview:** For Location Data
That wraps up our #PIIMasking1M 7 days announcement! HUGE thanks for following along and for your engagement.
Explore ALL our releases, including these E-PII previews, in the Ai4Privacy Hugging Face Collection & show some love ❤️ if you find them useful!
🔗 Visit the Collection:https://huggingface.co/ai4privacy
Let's keep building safer AI, together!
Standard PII tools often miss sensitive *business* data. That's why we built E-PII previews for the data that powers your operations and compliance needs.
Get a first look (representing 100,000 samples each!) into datasets designed for real-world enterprise security across these categories:
🏥 **PHI Preview**: For Healthcare Data
💳 **PFI Preview:** For Financial Data
🏢 **PWI Preview:** For Workplace Data
💻 **PDI Preview:** For Digital Activity Data
📍 **PLI Preview:** For Location Data
That wraps up our #PIIMasking1M 7 days announcement! HUGE thanks for following along and for your engagement.
Explore ALL our releases, including these E-PII previews, in the Ai4Privacy Hugging Face Collection & show some love ❤️ if you find them useful!
🔗 Visit the Collection:https://huggingface.co/ai4privacy
Let's keep building safer AI, together!
Post
1090
I need your help! Please vote for Ai4Privacy to do a demo at The first Global Open Source AI Conference for developers!
https://glosaic.org/demo-voting#:~:text=Founder%20@-,AI4Privacy
https://glosaic.org/demo-voting#:~:text=Founder%20@-,AI4Privacy
Post
2803
🚀 We are quite excited to announce the Ai4Privacy Python library! 🎉
pip install ai4privacy to anonymize short english text with OpenPII Masking 500k labels
📊 Day 5/7 of PII Masking 1M announcements complete! ⏰
pip install ai4privacy to anonymize short english text with OpenPII Masking 500k labels
📊 Day 5/7 of PII Masking 1M announcements complete! ⏰
Post
3077
🌟 Day 4: Two Models, One Privacy Mission! 🌟
The PII-Masking-1M series rolls on with two gems:
Categorical: ai4privacy/llama-ai4privacy-multilingual-categorical-anonymiser-openpii
Redaction: ai4privacy/llama-ai4privacy-multilingual-anonymiser-openpii
Join us in protecting data everywhere!
#AI #Privacy #OpenSource #Multilingual
The PII-Masking-1M series rolls on with two gems:
Categorical: ai4privacy/llama-ai4privacy-multilingual-categorical-anonymiser-openpii
Redaction: ai4privacy/llama-ai4privacy-multilingual-anonymiser-openpii
Join us in protecting data everywhere!
#AI #Privacy #OpenSource #Multilingual
Post
1739
📊 99%+ PII Masking Precision in English Straight to Your Browser! 🚀
ai4privacy/general-english-anonymiser-openpii-500k
Hard Facts:
🖥️ Runs in-browser—blazing fast, no server latency
👐 Open-source, MIT-licensed (even for commercial use)
📈 Full metrics on Hugging Face dataset and model pages
Day 3 out 7 of PII-Masking-1M Announcements Complete!
*Accuracies reported from the new OpenPII-500k dataset
#DataPrivacy #AI #OpenSource
ai4privacy/general-english-anonymiser-openpii-500k
Hard Facts:
🖥️ Runs in-browser—blazing fast, no server latency
👐 Open-source, MIT-licensed (even for commercial use)
📈 Full metrics on Hugging Face dataset and model pages
Day 3 out 7 of PII-Masking-1M Announcements Complete!
*Accuracies reported from the new OpenPII-500k dataset
#DataPrivacy #AI #OpenSource
Post
2126
#PII Masking Tech that does not **** around!
We are happy to release the OpenPII English Anonymiser —the most powerful open-source tool for redacting sensitive info from English text.
Fine-tuned Modernbert on 5.7 million+ PII examples, it’s clocking 99%+ accuracy across emails, dates, social numbers, and more!
Why it’s a big deal:
✅ Top-tier precision: 100% for passport numbers, 99.96% for emails*.
✅ Totally free: MIT license for personal or commercial use.
✅ No secrets: Full metrics shared on Hugging Face.
#AI #OpenSource #DataSecurity @huggingface
Day 2 out 7 of PII-Masking-1M Announcements Complete!
*Accuracies reported from the new OpenPII-500k dataset
ai4privacy/llama-ai4privacy-english-anonymiser-openpii
We are happy to release the OpenPII English Anonymiser —the most powerful open-source tool for redacting sensitive info from English text.
Fine-tuned Modernbert on 5.7 million+ PII examples, it’s clocking 99%+ accuracy across emails, dates, social numbers, and more!
Why it’s a big deal:
✅ Top-tier precision: 100% for passport numbers, 99.96% for emails*.
✅ Totally free: MIT license for personal or commercial use.
✅ No secrets: Full metrics shared on Hugging Face.
#AI #OpenSource #DataSecurity @huggingface
Day 2 out 7 of PII-Masking-1M Announcements Complete!
*Accuracies reported from the new OpenPII-500k dataset
ai4privacy/llama-ai4privacy-english-anonymiser-openpii
Post
2748
🚀 Ai4Privacy Team is excited to unveil PII-Masking-1M, our most significant release yet! 🎉
This publication series 📦 includes datasets 📊, models 🤖, and applications ⚙️ to advance PII masking with AI systems 🛡️
Starting on Monday with daily posts at 7 PM CET ⏰
This publication series 📦 includes datasets 📊, models 🤖, and applications ⚙️ to advance PII masking with AI systems 🛡️
Starting on Monday with daily posts at 7 PM CET ⏰