Modern enterprises are not only dealing with unprecedented data volumes, they’re also working within a more highly regulated environment than ever before. Maintaining security and compliance (while preserving usability) has become a high stakes game requiring the configuration and implementation of a number of different tools and processes.
Of these, data classification has quickly risen to the top of the pile. Let’s take a look at why that is, and how you can get started.
What is data classification?
Data classification is the process of organising data into relevant categories by identifying shared characteristics like sensitivity levels, risks, and applicable regulations. In Microsoft, this is achieved by applying sensitivity and retention labels and/or sensitive information type classifications. This can be done manually, or automatically using trained and trainable pattern-matching and classifiers.
Data classification makes it possible for organisations to apply appropriate protections at category level, rather than on an individual basis. Done right, it can dramatically reduce the time, effort and risk associated with modern data protection.
Why is data classification important?
To protect sensitive information appropriately, you need to know:
- What it is
- Where it is
- What value it holds
- What risk it presents
- Which regulations apply to it
- Who needs to have access
With a robust data classification process in place, all of these attributes can be tagged using labels and classifiers at document level, enabling those with similar protection requirements to be grouped for easier/more efficient handling.
The result is faster, more effective data protection that reduces security blind spots, lowers risk, eliminates redundancies, minimises storage requirements, and optimises resource allocation for more cost-effective IT.
What are the benefits of data classification?
The principal advantage of data classification is its ability to bring an organisation’s sensitive data to light. This is a critical step towards improving:
Data security
Being able to identify sensitive data types, locations and access requirements, and understand the specific implications of leaks/destruction/alteration, enables organisations to better:
- Understand the scope of data under their protection, and the variety of protections required.
- Consolidate sensitive information into appropriate storage to reduce its footprint and streamline data security.
- Reduce access to authorised users only, and ensure the right technologies are installed to prevent undesirable activity. (Encryption/data loss protection/information loss protection etc.)
- Optimise resource allocation to focus on high value data.
Regulatory compliance
Knowing what sensitive data is located where not only makes it easier to appropriately secure, it also makes it easier to trace and search. This enables organisation to:
- Handle different types of sensitive information in compliance with all applicable regulations.
- Find and retrieve specific data quickly to comply with data subject access requests.
- Prove appropriate compliance and security protocols are in place, and pass compliance audits.
Operational efficiency
By classifying data at point of creation (through established and/or automated data classification processes), organisations can protect, store and manage their data more effectively throughout its lifecycle. That means:
- More visibility into, and control over, data held and/or shared by the organisation.
- Easier (and safer) access to data for those with the appropriate authorisations.
- Greater insight into the value and risk associated with specific data, as well as the organisation’s data landscape as a whole.
- Enhanced retention and eDiscovery capabilities.
How does Microsoft Purview support data classification?
Microsoft supports data classification through the use of sensitivity and retention labels, as well as sensitive information type classifications.
Labels need to be preconfigured by admin to have a clear and sensible name, and perform a specific function (e.g. apply a ‘confidential’ watermark). Ideally, they should also include a tooltip to guide users in selecting the appropriate label when this is likely to be done manually.
Manual Classification
Manual classification requires users and/or admins to actively apply sensitivity or retention labels to content as they create and/or encounter it. Labels can be selected from the preconfigured list, or custom created.
Tooltips can be used to guide decision-making during this process, suggesting labels/classifications based on the contents of the document at hand.
Automated Classification
Sensitivity and retention labels can also be added automatically using automated pattern-matching that recognises:
- Specific keywords or metadata values
- Data matching known sensitive information patterns e.g. credit card numbers
- Content using specific templates
- Exact data matches
Pre-trained and trainable classifiers can also be used to tag content that isn’t easily identified by manual or automated pattern-matching. These classifiers focus more on what the entire item is than on the patterns contained within it. Common examples include legal and strategic business documents, pricing and other financial information.
Data Classification Dashboard
Once published, retention and sensitivity labels and automated classifiers need to be supervised to assess how effectively they’re being used, and what exactly is being done with them (and the data they apply to).
Microsoft supports administrators in this task via the Data Classification Dashboard. This provides a quick and intuitive overview of:
- how many items are classified under each sensitive information type
- top sensitivity labels in Microsoft 365 and Azure Information Protection
- top retention labels
- how users are interacting with sensitive content
- where sensitive and retained data is located
The dashboard also enables administrators to:
- configure and customise trainable classifiers
- configure and explore sensitive information types
- search for exact matches of a specific piece of sensitive data
- explore the contents of documents containing sensitive information (permissions dependent)
Where to start with data classification?
The best way to kick off any data classification journey is by bringing together key stakeholders for meaningful conversation around the challenges of data security and compliance, and the solutions Microsoft Purview offers.
Of course, that’s not always an easy task – particularly when you’re approaching the problem from a purely business/purely IT perspective. That’s where our compliance/risk experts come in.
With our combination of business, compliance and IT experience, we’re old hands at getting stakeholders from across the business on board, and facilitating effective discussions. Together, we can ensure a pragmatic and fit-for-purpose data classification taxonomy is implemented. Talk to us today about how we can help.