Publication: Promptshield: policy-aware data loss prevention framework for generative AI prompts
Loading...
Date
Authors
Advisor
Department
Computer Engineering
Journal Title
Journal ISSN
Volume Title
Publisher
Graduate School
Type
Abstract
Generative artificial intelligence (GenAI) systems that are capable of generating content, summarizing documents, analyzing trends, assisting with coding, and automating communications have altered the nature of knowledge work. However, this shift creates new security and privacy issues. Employees use free-form prompts when using large language models (LLMs), which often include personal identifiable information, financial data, proprietary research materials, and/or strategic corporate insights. Unlike typical enterprise software systems that operate in defined parameters, GenAI accept unrestricted text from users to provide a new, largely unprotected, data-exposure point. When users enter their prompts into cloud-based LLMs, the prompts are recorded, stored, used to enhance the model(s), and/or stored in regions over which organizations have no jurisdiction. These risks expose a significant weakness in current Data Loss Prevention (DLP) architectures. Present-day DLP architectures are focused primarily on monitoring email traffic, file uploads/downloads, data storage repositories, and/or user endpoints. Traditional DLP architectures focus on structured data flows and static content. Therefore, traditional DLP architectures are ineffective against the ephemeral, conversational nature of GenAI prompts residing temporarily in the interface before being sent to an Application Programming Interface (API). Furthermore, safety filters provided by LLM vendors are generally limited to controlling what a model produces after it receives the user's input, and do not prevent sensitive information from being communicated to an LLM in the first place. The objective of this dissertation is to close this gap by providing a Client Side, Policy-Aware, Real Time Data Loss Prevention (DLP) Framework (PromptShield) that is designed to protect GenAI prompts from users entering sensitive information. PromptShield functions as an interceptive layer that examines each user input prior to submission to an LLM. To accomplish this goal, PromptShield was developed to perform lightweight computations with a low-latency response while seamlessly integrating into existing enterprise workflows. The PromptShield framework is comprised of three interconnected components which are semantic classifier, a role-based policy engine, and an entity-aware redaction component. The semantic classifier utilizes a combination of MiniLM Sentence Embeddings and a Random Forest (RF) Model to classify user prompts as either Safe, Risky, or Confidential. The output of the classifier provides a baseline risk assessment, which is further enriched by the role-based policy engine. The role-based policy engine takes the classification produced by the semantic classifier and maps it to a user-defined Allow, Warn, or Block action based upon the user's organizational role. The entity-aware redaction component of the system will enable users to anonymize sensitive entities that are identified by the system as part of the Named Entity Recognition (NER) process. The combination of this entity-aware redaction capability along with the other described capabilities provide for a trade-off between providing high levels of protection for the data, while minimizing disruption in the user's workflow. For the purposes of training the dataset consisted of 900 generated enterprise prompts spread evenly across the domains of Human Resources, Finance, and Research \& Development. Each prompt included a sensitivity label. A test set of 435 unseen prompts were utilized to evaluate the generalizability of the classifier. The classifier performed with an average accuracy of 72.92\%. Compared to large transformer models, the accuracy may be slightly lower; however, it is suitable for a real-time client-side application optimized for speed and response. Examination of the confusion matrix shows that the majority of misclassifications occur between Safe and Risky categories. Misclassifications within the borderline categories, e.g., "Can I review the salary distribution data for trend analysis?" is not directly indicative of HR-related confidential information, yet does reference a potential source of HR-confidential information indirectly. Contextual subtleties, such as these, are difficult for the lightweight embedding models to recognize. However, the classifier demonstrated a very high recall rate for the Confidential category, thus minimizing the risk of high-risk false-negatives. The role-based policy engine demonstrated its capability to dynamically adjust the risk assessments based on the role of the user. In particular, HR-related prompts were frequently classified as Warning because they contained personal identifiers and internal employee data. Conversely, Finance-related prompts had a much more relaxed pattern, reflecting the organization-wide norms governing the handling of transactional/budgetary information with fewer constraints. Meanwhile, R\&D-related prompts showed the highest percentage of Block actions, consistent with the strictest confidentiality standards of research domains. This variation supports the notion that DLP enforcement must be context-aware, i.e., the same sensitivity classification can represent different levels of risk depending on the department utilizing the LLM. The redaction module was evaluated qualitatively and demonstrated the ability to effectively mask sensitive information while maintaining the integrity of the original sentence structure. The redaction module relied on Named Entity Recognition (NER) to identify various types of high-risk expressions common to enterprise prompts, including PERSON, ORG, Geo-Political Entity (GPE), MONEY, DATE, and Location (LOC). These Redacted Outputs retained structural coherence so that users were able to modify their prompt without losing its intended meaning. One of the major usability issues that are inherent in traditional DLP solutions is that they completely block all of the content and the user either becomes frustrated or will find ways to circumvent it. With the ability for users to safely reformulate their original content, PromptShield allows organizations to protect their data by maintaining productivity. To demonstrate the practicality of implementing PromptShield in real world business environments, the entire PromptShield pipeline was evaluated to determine processing time. The total processing time for the PromptShield pipeline (embedding, classification, policy mapping and optionally redacting) was approximately 120-150 milliseconds per prompt when run on a standard laptop with no GPU acceleration. These results indicate that the PromptShield System could be run on user's local machines without negatively impacting perceived performance from the user; therefore it is suitable for various use cases including browser add-ons, chat interfaces, Integrated Development Environments (IDE), and Corporate Communication Platforms. The discussion of results highlights that although PromptShield performs well in all key areas of evaluation, there are still several challenges to overcome. The most significant limitation is detecting sensitivity in prompts where risk is inferred through implied context rather than explicit entities. Higher precision in these contexts may be achievable through the use of domain-specific embeddings, or through hybrid models that utilize metadata during inference. A second limitation of PromptShield is its reliance on general-purpose NER. While general-purpose NER is effective for identifying common entity types, it may fail to identify organization-specific identifiers, such as internal project codes, proprietary acronyms, or confidential system labels. Future work may include fine-tuning NER models on domain-specific datasets. Also, Adaptive learning from user feedback, behavior based analytics or risk scoring will be important in developing a more responsive and dynamic enforcement mechanism for PromptShield in the future since current policy is static. In spite of these limitations, the results indicate that prompt-level DLP is technically feasible and strategically necessary for modern organizations. PromptShield closes a significant gap created by traditional cybersecurity tools and LLM vendor safety mechanisms by focusing on the input stage of human and Artificial Intelligence (AI) interactions. PromptShield presents a unique, practical, and effective way to prevent unintended data leakage in environments where employees increasingly rely on GenAI systems. The design of PromptShield makes it easy to integrate into production environments due to its light-weight design, modular architecture, and ease of integration. Ultimately, the thesis demonstrates how PromptShield fills an urgent need for enterprises, while establishing a solid base for developing the next generation of AI governance, responsible AI deployment, and data protection enterprise models. The work is part of a large body of research which emphasizes the need for transparency, user-centered development, and real-time control with regard to AI systems. With continued use of GenAI by companies, as such as PromptShield, frameworks will be increasingly important in protecting confidentially, assuring compliance, and managing ethically data while providing productivity.
Description
Thesis (M.Sc.) -- Istanbul Technical University, Graduate School, 2026
Journal or Series
ISSN
ISBN
Rights
Keywords
Cyber security, Siber güvenlik, Artificial intelligence, Yapay zeka
Citation
Collections
Endorsement
Review
Supplemented By
Referenced By
34
Görüntülenme
149
İndirme
Google Scholar
Scholar'da Ara ↗ Bu yayında DOI yok — Altmetric/Dimensions/PlumX/BIP! rozetleri DOI gerektirir.