MDDI スピーチ · 2025-07-07

Personal Data Protection Week 2025におけるMinister Josephine Teoの開会挨拶

楊莉明 · デジタル開発・ニュース相 · 個人データ保護週間

要点

  • シンガポールのIMDAは8か国と共同で初の地域レッドチーミング・チャレンジを開催し、大規模言語モデルにおけるステレオタイプに基づく出力——特定の民族的な名前を犯罪的役割と結びつけるものなど——を明らかにしました。これらは学習データに内在するバイアスに起因するものです。
  • 3年間にわたり、IMDAとPDPCはPET Sandboxを運営し、Ant Internationalをはじめとする企業が、生の顧客データを交換することなく、Privacy Enhancing Technologiesを通じてパートナーと共同でAIモデルを訓練できる環境を提供してきました。その結果、バウチャー引き換え率において測定可能な改善が得られています。
  • IMDAは、経営幹部(C-suite)を対象としたPETs Adoption Guideを公開する予定です。これは、各組織がそれぞれのビジネスニーズに適したPrivacy Enhancing Technologiesを特定し、導入にあたっての重要な考慮事項を把握できるよう支援するものです。
  • IMDA、AI Verify Foundation、および業界パートナーは、生成AIアプリケーションの信頼性をテストするための標準化された手法を開発するため、Global AI Assuranceパイロットを実施しました。その成果は、望ましくないコンテンツや意図しないデータ漏洩といったリスクをカバーする「IMDA Starter Kit」に集約されています。
  • IMDAは、このパイロットを新たな継続的な取り組みであるAI Assurance Sandboxへと移行させています。同Sandboxでは、ビジネスユーザー、ガバナンスチーム、およびAI開発者が協力して、生成AIアプリケーション向けのガードレールとテストプロセスを共同で開発することができます。
  • IMDAはEnterprise SingaporeおよびSingapore Accreditation Councilと連携し、Data Protection Trustmarkを新たな国家標準——Singapore Standard 714——へと格上げしました。これにより、データ保護における卓越性を証明しようとする組織に対して、正式な認証ベンチマークが提供されます。

全文翻訳

MDDI 英語原文の翻訳 · 翻訳日: 2026-06-21

皆さん、おはようございます。まず、本日ご参集いただいた全ての皆様に感謝申し上げます。本日は会場に1,500名以上がお集まりで、アジア各国やさらに遠方からお越しの方々を含め、今週を通じて2,000名を超える方々にご来場いただく予定です。特に、ASEAN加盟国のデータ保護当局をはじめとする国際的なゲストの皆様のご参加を、心より歓迎申し上げます。ご出席いただいた全ての皆様、ありがとうございます。

今年のテーマは「変化する世界におけるデータ保護」です。これは、グローバルな事業環境と技術の世界の両方において生じている大きな変化を認識したものです。

これら二つの力は、私たちの職場、家庭、そして相互の関係を変容させています。私たちは自らの慣行、法律、さらには広範な社会規範をも見直していくことが避けられません。

ここにお集まりの皆様の多くは、データまたはAI、あるいはその両方の実務家であると思います。

昨年、私はAI時代におけるデータの重要性についてお話ししました。この点は今もなお非常に重要です。生成AIモデルが膨大な量のデータの上に構築されており、事前学習からファインチューニング、テストおよび検証に至るAI開発ライフサイクル全体を通じて、データが不可欠であることは皆様もよくご存知のとおりです。

近年、カスタマイズされたまたは独自のデータセットを基盤とした、分野特化型AIアプリケーションが急増しています。

良い例として、乗客からの問い合わせに対応するためのチャンギ空港のチャットボット「AskMax」が挙げられます。これはチャンギ空港のデータリポジトリを呼び出すよう設計されたLLMの上で動作しています。

もう一つの例は「GPT-Legal」であり、これはIMDAがSingapore Academy of LawのLawNetデータベースを使用してファインチューニングしたものです。

AI時代におけるデータの重要性を考えれば、データが継続的な進歩を制約する要因ともなっているのは驚くことではありません。

AI開発と活用の各段階におけるデータの課題を順に見ていきましょう。

モデルの学習段階において、まず広く知られている問題は、大規模モデルの学習にインターネット上のデータが使用されていることです。インターネット上のデータは品質にばらつきがあります。また、ディスカッションフォーラム上のユーザー生成コンテンツをはじめとするさまざまなソースから、偏ったまたは有害なコンテンツが含まれていることも多いです。基礎となるデータ入力に有害・毒性・偏向のあるコンテンツが含まれている場合、モデルの出力に下流の問題を引き起こす可能性があります。

シンガポールのIMDAと他の8か国が共同で実施した初の地域レッドチーミング・チャレンジにおいて、問題のあるモデルの挙動が観察されました。シンガポールの受刑者に関するスクリプトを書くよう求められた際、LLMは違法賭博で収監されたキャラクターに「Kok Wei」、酔っぱらいによる迷惑行為者に「Siva」、薬物乱用違反者に「Razif」といった名前を選びました。学習データから取り込まれたと思われるこれらのステレオタイプは、まさに私たちが回避すべきものです。

同時に、開発者はインターネット上のデータが枯渇しつつあるという問題に直面しています。ほとんどのLLMはすでにインターネット上のデータのコーパス全体を学習済みです。では、モデル提供者はモデルをどのように改善すればよいのでしょうか。彼らはモデルを補強するために、より機密性が高くプライベートなデータベースに目を向けており、これ自体が新たな一連の課題をもたらしています。

例えばOpenAIは、グローバルなニュースメディアだけでなく、アイスランド政府、Apple、Sanofi、アリゾナ州立大学など、政府、企業、大学とのデータ関連パートナーシップを拡大し続けています。

パートナーシップモデルはデータの可用性を高める一つの方法ですが、時間がかかりスケールアップが困難です。これらのデータベースの一部には、個人データや事業上の機密情報などの機密データが含まれている可能性があります。

機密情報を保護しながらモデルを学習させる手段が、ますます求められています。

AIモデルの上に重ねられた「スキン」とも言えるAIアプリケーション(「アプリ」)も、信頼性上の懸念をもたらす可能性があります。アプリが不正確・偏向・有害な情報を提供したり、機密情報を漏洩したりした場合、企業の評判に深刻な影響を及ぼし、最悪の場合には実際に身体的被害を引き起こすこともあり得ます。

通常、企業はアプリの信頼性を確保するために、さまざまな既知のガードレールを採用しています。具体的には、モデルの挙動を誘導するための詳細なシステムプロンプトの作成、正確性を向上させるための検索拡張生成(RAG)の活用(皆様もよくご存知かと思います)、機密情報をふるい分けるためのさまざまな種類のフィルターなどが含まれます。

それでもなお、アプリには予期せぬ欠点が生じることがあります。サードパーティのテスター「Vulcan」は最近、見込み顧客から寄せられる製品仕様に関する質問への回答を従業員が支援するためのある高技術メーカーのチャットボットをテストしました。そのメーカーは、アプリが誤って機密ビジネス情報——例えば、見込み顧客に知らせたくない事柄——を漏洩してしまうことを懸念していました。果たして、Vulcanは中国語でプロンプトを入力した際に、アプリがバックエンドの販売コミッション率を漏洩することを発見しました。メーカーの立場から考えると、見込み顧客に販売コミッション率を伝えることは、値引きの余地がどれほどあるかを明かすことに等しく、どのような企業にとっても望ましいことではありません。

幸い、この問題はテスト段階で発見されました。これは独立したテストの価値を示しています。リリース前にGenAIアプリの信頼性を確保するためには、アプリが意図どおりに機能しているか、また安全性の一定のベースラインが確保されているかを、体系的かつ一貫した方法で確認することが重要です。

モデル開発者と同様に、アプリ開発者もデータの不足という問題に対処しなければなりません。多くの場合、モデルは企業の特定のニーズに対応するために社内データベースと連携していますが、信頼性の高いアプリを構築するための独自データが不足していることが多いです。IBMのグローバル調査への回答者の42%が、これをAI導入における最大の課題の一つとして挙げています。したがって、機密情報を保護しながら、企業間でより多くのデータ共有を促進する手段が必要です。

AIアプリが展開され消費者に使用された後、誤った情報や有害な情報を修正することは大きな課題となります。モデルが何かを「学習」した後に行うファインチューニングや再学習のプロセスは不正確であることが多く、コストも高くなりがちです。

そのため、「マシン・アンラーニング」(機械的忘却)は新たな分野として台頭しつつあります——ただし、まだ萌芽的な段階にあります。Anthropicのような主要なLLMリーダー企業が直面している主要な課題は、現在のモデルが数十億から数兆ものパラメーターを持つという点です。どの変数が出力の欠点に最も寄与しているのでしょうか。それらを特定し、大規模で的を絞ったモデル修正を実施するための技術は存在するのでしょうか。

最後に、説明責任という根本的な課題があります。AIのライフサイクルは、モデル構築者、デプロイヤー、ユーザーなど多くの関係者が関与しており、複雑です。それぞれがリスクを軽減する役割を担っています。

ここにお集まりのコミュニティの皆様は、Samsungの従業員グループがエラーチェックのために機密のソースコードをChatGPTに貼り付けることで、意図せず機密情報を漏洩してしまった事例をご存知かと思います。これは私たちの職場でも起きているとも思います——同僚がスペルチェックや文章表現の確認のために、ファイルをChatGPTにアップロードすることがあるかもしれません。そのファイルの中に、ChatGPTと共有すべきでない情報が含まれていないか、考えさせられます。

機密情報をチャットボットに入力してはいけなかった従業員の責任でしょうか?ここにいる同僚のほとんどは、彼らにある程度の責任があると考えているでしょう。

しかし、機密データが収集されないよう十分なガードレールを確保することは、アプリ提供者の責任でもあるのでしょうか?

あるいは、そのようなデータがさらなる学習に使用されないよう確保することは、モデル開発者の責任であるべきでしょうか?

残念ながら、これらに対する容易な答えはありません。

AIが進歩し続けるためには、組織的なプロセスの改善から新たなリスク軽減技術の開発まで、さまざまな種類のソリューションが必要です。プライバシーを損なうことなくデータの活用を最適化するプライバシー強化技術(PETs)などの技術的ソリューションが、これらの懸念に対処するための実行可能な経路として浮上しています。

過去3年間、IMDA および PDPC は、さまざまなセクターおよびユースケースにおけるプライバシー強化技術(PETs)の活用を企業が探索・実験できるよう、PET Sandbox を運営してきました。関心は高まりを見せており、アーリーアダプターの一部は具体的なビジネス上の成果も経験しています。

例えば、Sandbox に参加した金融機関である Ant International は、複数の PETs を組み合わせ、顧客情報を互いに開示することなく、デジタルウォレットパートナーとともに AI モデルを訓練しました。その目的は、ウォレットパートナーが提供するバウチャーを、最も利用する可能性の高い Ant International の顧客にマッチングするモデルの構築にありました。Ant International は顧客のバウチャー利用データを提供し、デジタルウォレット企業は同一顧客の購買履歴・嗜好・デモグラフィックデータを提供しました。AI モデルはそれぞれのデータセットで個別に訓練され、データオーナーが互いのデータを閲覧・取り込むことはありませんでした。この結果、バウチャー取得件数が大幅に向上し、ウォレットパートナーの収益増加、そして Ant International の顧客エンゲージメント向上につながりました。

このような PETs の活用方法には、例えば不正検知や、医療機関が患者をより適切にケアするための活用など、多くのユースケースがあることがおわかりいただけるかと思います。

合成データ(Synthetic Data)は、大きな可能性を示す PET のもう一つの例です。昨年、私は PDPC の「合成データ生成ガイド」を発表し、組織向けのベストプラクティスを示しました。現在、シンガポールには Betterdata のような革新的な企業が登場し、AI 開発者が実世界のデータセットを模倣したデータを生成できるよう支援しています。こうした合成データは、AI モデルを構築するための学習データセットとして既存のデータセットをさらに拡充するものであり、先ほど言及したデータに関する課題の解決に一定程度貢献しています。

Sandbox に参加した組織との経験を通じて、私たちはこれらの技術、およびデータが共有される際の個人データ保護や法的義務への対応能力をより深く理解することができました。また、PET ソリューションの提供に関心を持つ技術プロバイダーや、PETs の活用に積極的な企業からの関心の高まりについても、確かな感触を得ています。

この勢いをさらに発展させるべく、IMDA は PETs Adoption Guide を導入する予定です。経営幹部(C スイート)を対象に設計されたこのガイドは、組織がビジネスニーズに適した PETs を特定するためのリソースを提供するとともに、PETs を効果的に導入するための主要な考慮事項も含んでいます。

今年の Personal Data Protection Week でも、PETs Summit が引き続き開催されます。昨年、初めて開催された際と同様に、今回のサミットも、データ保護当局、既存および参加を検討している PETs ソリューションプロバイダー、そして Sandbox の利用者が互いに交流し、学び合う貴重な機会となるでしょう。

PETs Sandbox で実証されているように、シンガポールの新興技術に対するアプローチは、企業が実験できるツール・リソース・安全な環境を提供し、その学びを迅速に共有することで産業と消費者が恩恵を受けられるようにするというものです。

最近、IMDA、AI Verify Foundation および業界パートナーは、Global AI Assurance パイロットに共同で取り組み、生成 AI アプリケーションの信頼性を検証する方法を研究しました。テストは、AI アプリケーションが主要なリスクに対処していることを証明するための重要なステップです。

家庭の家電や職場への移動手段となる車両など、私たちが日常的に使用するものの多くは、適切にテストされていなければ使用しないでしょう。しかし現状では、AI アプリケーションが適切なテストを経ないまま日常的に私たちに対して使用されています。これはまさに空白、すなわち埋めなければならない深刻なギャップです。

一例として、Changi General Hospital があります。同病院はサードパーティテスターである Softserve と連携し、特定の医療報告書の要約ツールの信頼性をテストしました。他の医師と共有できる症例・患者サマリーを作成できることは、医師とその業務量の観点から非常に有益です。この要約ツールの信頼性・正確性を確保し、患者情報が誤って伝わらないようにすることは、最重要課題です。

もう一例は NCS であり、同社はコーディングアシスタントが社内のコーディング標準およびセキュリティ要件、ならびに外部の規制ガイドラインにどの程度準拠しているかをテストしました。

このパイロットから得られた知見をもとに、IMDA は組織がリスクをテストし管理するために活用できるいくつかのテスト手法を特定しました。このテスト手法の集成は「IMDA Starter Kit」として知られています。これは、ガバナンスフレームワークやガイドラインを超えた、AI アプリケーションのテストと展開に関するより標準化された手法を求める企業の声に直接応えるものです。先ほど説明した望ましくないコンテンツや意図しないデータ漏洩といったリスクへのテストも含まれています。

IMDA がパイロットを新たな継続的 AI Assurance Sandbox に移行するにあたり、学習と改善のサイクルは続いています。この Sandbox は、ビジネスユーザー、ガバナンスチーム、AI 開発者を問わず、生成 AI アプリケーションに対するより優れたガードレールやプロセスといったソリューションを共同で開発するための学習環境です。自社アプリケーションをテストに供し、共有知識基盤に貢献することに関心のある組織の参加を歓迎します。

最終的に、これらの各 Sandbox における私たちの目標は、データ保護であれ AI ガバナンスであれ、「良い状態」とはどのようなものかについて連携と合意を形成することです。

製品安全性や医薬品などの伝統的な分野と同様に、維持すべき標準について専門家が合意し、その標準が満たされていることを保証するテスターが必要です。

AI 導入の速度と規模を考慮すると、標準の策定と合意には一定の緊急性があります。現実的には、これには時間がかかるでしょう。経るべき段階が数多くあります。少なくともシンガポールでは、テストおよびアシュアランスのエコシステムを育成するための重要な最初のステップを踏み出しています。業界プレーヤーが私たちとともに、最終的な正式標準の確立の基盤となり得る「ソフト」標準を先行して形成してくださることを願っています。

データ保護の分野は先行しており、次のステップへの準備が整ったことを喜んでお伝えします。

IMDA は Enterprise SG およびシンガポール認定協議会(Singapore Accreditation Council)と連携し、データ保護トラストマーク(DPTM)を新たなシンガポール規格である Singapore Standard 714 に格上げしました。説明責任あるデータ保護の実践を実証した企業は、この新規格の下で認証を申請することができ、データ保護の卓越性を示したい企業にとっての国家基準が設定されます。このトラストマークは、認証を受けた組織が個人データの保護においてワールドクラスの実践を採用していることを消費者に保証するものです。

AI の発展のためにデータを活用する際の課題と機会に対するシンガポールのアプローチについて、ご理解いただけたと思います。

AI が責任を持って開発され、信頼性をもって展開される場合、データを解放するための手法も含め、企業と人々が得られるものは多いと信じています。いかにそれを実現するかを理解し、適切な措置を講じることは、企業および政府のリーダーとしての私たち次第です。

そうすることで、AI の導入を促進するのみならず、データおよび AI ガバナンスへの信頼をさらに高めることができるでしょう。このことを念頭に置き、今後の議論が実りあるものとなることを願っています。誠にありがとうございました。

英語原文

MDDI 公式サイト原文 · 取得日: 2026-06-21

Good morning, colleagues and friends. I’d first like to thank everyone for being here. We have over 1,500 people in the room today, and over 2,000 coming and going throughout the week, including from many countries in Asia, and even further afield. I especially appreciate our international guests for joining us, including Data Protection Authorities from fellow ASEAN member states. Thank you all for being here.

The theme for this year is “data protection in a changing world”. This is an acknowledgement of the significant changes in both our global operating environment, as well as in the world of technology.

These twin forces have disrupted our workplaces, our homes, and our relationships with each other. It is inevitable that we must adjust our practices, laws and even our broader social norms.

Most of you in this room are practitioners of data or AI, or both.

Last year, I had spoken about the importance of data in the age of AI. This remains as pertinent as ever. We all know that generative AI models are built on vast amounts of data, and data is critical throughout the AI development lifecycle, from pre-training, to fine-tuning, to testing and validation.

In recent times, we have seen an explosion of sector-specific AI applications built on customised or proprietary datasets.

A good example is AskMax, Changi Airport’s chatbot that helps to address passenger queries. It runs on a LLM designed to call on Changi Airport’s data repositories.

Another example is GPT-Legal, which was finetuned by IMDA using the Singapore Academy of Law’s LawNet database.

Given the criticality of data in the AI age, it should not be surprising that data has also become a limiting factor to continuing advancement.

Let us walk through the data challenges at each stage of AI development and use.

In model training, the first well-known issue is the use of internet data to train these large models. Internet data is uneven in quality. Often, they contain biased or toxic content from different sources, including user-generated content on discussion forums. When the underlying data input contains harmful, toxic or biased content, this can lead to downstream problems with model outputs.

In the first regional red teaming challenge run jointly by Singapore IMDA and eight other countries, problematic model behaviours were observed. When asked to write a script about Singaporean inmates, the LLM chose names such as “Kok Wei” for a character jailed for illegal gambling, “Siva” for disorderly drunk and “Razif” for a drug abuse offender. These stereotypes, most likely picked up from the training data, are actually things that we want to avoid.

At the same time, developers are running out of internet data. Most of the LLMs are already trained on the entire corpus of internet data. What then should model providers do to improve their models? They are turning to more sensitive and private databases to augment their models, which brings its own set of challenges.

OpenAI, for example, has a growing list of data-related partnerships not only with global news outlets, but also governments, companies and universities like the Icelandic Government, Apple, Sanofi and Arizona State University.

The partnership model is one way of increasing data availability, but it is time-consuming and difficult to scale. Some of these databases may include sensitive data such as personal data or business confidential information.

Increasingly, we need a way to train models, while protecting sensitive information.

AI application, or ‘app’, which can be seen as the ‘skin’ that is layered on top of AI models, can also pose reliability concerns. If apps provide inaccurate, bias or toxic information, or leak confidential information, these can have serious implications for the company’s reputation, and in the worst cases, may actually cause physical harm.

Typically, companies would employ a range of well-known guardrails to make their app reliable. These include writing detailed system prompts to steer the model behaviour, using retrieval-augmented generation (or RAG), which many of you are familiar with, to improve accuracy or different types of filters to sieve out sensitive information.

Even then, apps can have unexpected shortcomings. Vulcan, a third-party tester, recently tested a high-tech manufacturer’s chatbot that assists employees to answer questions on product specifications that are posed by prospective customers. The manufacturer was concerned that the app would inadvertently leak confidential business information, for example, telling the prospective customers something that they do not want the prospective customers to know. True enough, Vulcan found that when prompted in Mandarin, the app leaked backend sales commission rates. You can imagine, from the manufacturer’s point of view, telling the prospective customers what the sales commission rates are is basically revealing how much further they can cut the price – and it is not something any business wants.

Fortunately, this problem was discovered during the testing phase. This highlights the value of independent testing. To ensure the reliability of GenAI apps before release, it is important to have a systematic and consistent way to check that the app is functioning as intended, and there is some baseline safety.

Like model developers, app developers must deal with data inadequacies. Very often, the models are linked up with internal company databases so that the apps can cater to the businesses’ specific needs. However, there are often insufficient proprietary data to build reliable apps. 42% of respondents to an IBM global survey cited this as one of their biggest challenges to AI adoption. So, we need a way to unlock more data-sharing among companies while protecting sensitive information.

After AI apps are deployed and used by consumers, correcting erroneous or harmful information poses a significant challenge. The process of finetuning and retraining a model – after it has “learnt” something – is imprecise and often costly.

Machine unlearning has therefore become a new field, albeit a nascent one. A key challenge faced by LLM leaders like Anthropic is that models now have billions or trillions of parameters. Which variables contribute most to the shortcomings in output? Are there techniques to identify them and carry out targeted model corrections at scale?

Finally, an overriding concern is accountability. The AI lifecycle is complex, with model builders, deployers, users and more. Each has a role to play to mitigate the risks.

This community here would be familiar with the case of a group of Samsung employees who unintentionally leaked sensitive information by pasting confidential source code into ChatGPT to check for errors. I think we are aware that this is happening in our workplaces too – sometimes our colleagues, in order to do a spell check, or to check the way in which they have put across ideas, may upload a file on to ChatGPT. This makes you wonder if there is anything in the file that should not be shared with ChatGPT.

Is it the responsibility of the employees who should not have put sensitive information into the chatbot? I think most of our colleagues here believe they have some responsibility.

But is it also the responsibility of the app provider to ensure that they have sufficient guardrails to prevent sensitive data from being collected?

Or should model developers be responsible for ensuring that such data is not used for further training?

There are no easy answers to this, I’m afraid.

For AI to continue advancing, we will need various types of solutions – from organisational process improvements to developing new techniques in risk mitigation. Technical solutions, such as Privacy Enhancing Technologies – or PETs that optimise the use of data without compromising privacy – have emerged as a viable pathway for addressing these concerns.

In the last 3 years, the IMDA and PDPC have run the PET Sandbox to encourage businesses to explore and experiment with the use of PETs across a variety of sectors and use cases. We have seen growing interest and some early adopters have also experienced tangible business returns.

For instance, Ant International, a financial institution that joined the Sandbox, used a combination of different PETs to train an AI model with their digital wallet partner without disclosing customer information to each other. The intention was to use the model to match vouchers offered by the wallet partner with customers of Ant International, who were most likely to use them. Ant International contributed voucher redemption data of their customers, while the digital wallet company contributed purchase history, preference and demographic data of the same customers. The AI model was trained separately with both datasets, without each data owner seeing or ingesting the other’s data. This led to a vast improvement in the number of vouchers claimed; the wallet partner increased its revenues, while Ant International enhanced its customer engagement.

You can see that this way of using PETs has many use cases, for example in detecting fraud, or in allowing healthcare institutions to do a better job of taking care of their patients.

Synthetic Data is another example of a PET that shows good promise. Last year, I launched PDPC’s Guide on Synthetic Data Generation, which sets out best practices for organisations. There are now innovative companies in Singapore, such as Betterdata, that help AI developers generate data to mimic real-world datasets. These synthetic data can further augment existing datasets as training datasets to build AI models, which goes some way to addressing the data challenges I had referred to earlier.

Our experience with organisations in the Sandbox has allowed us to better understand the technologies, their ability to protect personal data and comply with legal obligations when such data is shared. It has also given us a good sense of the growing interest from technology providers in offering PET solutions, as well as companies who are keen to use PETs.

To build on this momentum, IMDA will be introducing a PETs Adoption Guide. Designed for C-suite executives, this guide will offer resources to help organisations identify the right PETs for their business needs and will also include key considerations for companies to effectively deploy PETs.

This year’s Personal Data Protection Week will once again include the PETs Summit. Similar to last year when it was held for the first time, the Summit will be a good opportunity for data protection authorities, existing and interested PETs solution providers, and users in the Sandbox to connect and learn more from one another.

As demonstrated in the PETs Sandbox, Singapore’s approach towards emerging technologies is to help provide tools, resources, and a safe environment for companies to experiment, and to quickly share the learnings so that industries and consumers can benefit.

Recently, IMDA, AI Verify Foundation and industry partners collaborated on a Global AI Assurance pilot, studying ways to test the reliability of generative AI applications. Testing is a critical step to demonstrate that the AI application has addressed key risks.

A lot of the things that we use on a day-to-day basis, such as the appliances in our homes, the vehicles that take us to the workplace – we would not use them if they had not been properly tested. And yet, on a day-to-day basis, AI applications are being used on us without having been properly tested. So this is a lacuna, a serious gap that needs to be filled.

One example is Changi General Hospital, which worked with third party tester Softserve to test the reliability of their summarisation tool for selected medical reports. It is incredibly helpful to doctors and their workloads, to be able to put together case or patient summaries that can be shared with other physicians. How we ensure that this summarisation tool is reliable, accurate and does not misrepresent the patient, is of utmost importance.

Another is NCS, which tested how well its coding assistant adhered to internal coding standards and security requirements, as well as external regulatory guidelines.

With insights from this pilot, IMDA has identified several testing methods that organisations can use to test for and manage risks. This compilation of testing methods is known as the “IMDA Starter Kit”. It is a direct response to companies’ requests to go beyond governance frameworks and guidelines, for more standardised ways to test and deploy AI applications. It includes testing for risks like undesirable content and unintended data disclosure, like those I described earlier.

The learning and iterating continue as IMDA transitions its pilot to a new, ongoing AI Assurance Sandbox. The Sandbox is a learning environment to help all of us – whether we are business users, governance teams, AI developers – to jointly develop solutions, like better guardrails or processes for gen AI applications. Organisations interested in putting their applications to the test and contributing to our shared knowledge base are welcome to join.

Ultimately, our aim with each of these Sandboxes is to find coalition and consensus around what good looks like, whether for data protection or AI governance.

Much like traditional fields of product safety or pharmaceuticals, we need subject matter experts to agree on the standards to uphold, and testers to assure us that the standards are being met.

Given the speed and scale of AI adoption, there is some urgency for standards to be developed and agreed to. Realistically, this will take time. There are many stages to go through. In Singapore at least, we have taken the critical first steps to grow the ecosystem for testing and assurance. Our hope is that industry players will join us to initiate ‘soft’ standards that can be the basis for the eventual establishment of formal standards.

The field of data protection has had a head start, and I am pleased to share that we are ready to take the next step.

IMDA has worked with Enterprise SG and the Singapore Accreditation Council to elevate the Data Protection Trustmark (DPTM) to a new Singapore Standard, Singapore Standard 714. Companies that demonstrate accountable data protection practices can now apply to be certified under this new Standard, which will set the national benchmark for companies that want to demonstrate data protection excellence. The Trustmark will assure consumers that certified organisations adopt world-class practices in protecting their personal data.

I hope I have given you a sense of Singapore’s approach to dealing with the challenges and opportunities in using data for AI advancement.

We believe there is much for businesses and people to gain when AI is developed responsibly and deployed reliably, including the methods for unlocking data. It is up to us as leaders in corporations and the government to understand how we can do so, and to put in place the right measures.

By doing so, not only will we facilitate AI adoption, we will also inspire greater confidence in data and AI governance. On that note, I wish you fruitful discussions in the days ahead. Thank you very much.