Data Sharing Agreement as part of data access rights management
Hello, we are looking into the end-to-end process for authorizing access to data in our main data lake and access to reports built from this data. One of the tools we could leverage is the Data Sharing Agreement. But before we jump to finding answers with Collibra, we have to ask the good questions. What kind of rules should drive data access? If you have started this side of data governance in your perimeter, please share your thoughts and ideas.
marieaudemagarshack
Posted 4 years ago · Edited 1 year ago·Last reply 4 years ago
17 comments
marieaudemagarshack
OP4 years ago · EditedHello David, on our side, yes, we intend to relate a data sharing agreement to the purposes for sharing the data. Depending on the purposes picked by the consumer requesting access to the data, different rules or standards may apply. We are probably going to prepare a pre-defined list of purposes, as we did to build a GDPR Processing Register. Hope this helps.
davidbotzenhart
·4 years ago · EditedThis is such a helpful thread! Questions for the community, along with the data sharing agreement do you also incorporating a usage purpose asset? The purpose asset is a standard or typical purpose of data usage that the governance team maintains. Such as power bi report development or data science experiment. This would be selected in the data usage request and helps in keeping a consistent response to the what are you using this data for.
marieaudemagarshack
OP4 years ago · EditedHello Brad, thank you for this detailed reply. From our end, we set confidentiality, personally identifiable information (PII) and sensitive personal information on data elements., using the same 4 confidentiality levels: public, internal, restricted, confidential. We start from the assumptions that
As we are maturing this topic, we are also considering additional criteria that we may want to take into account in the data sharing agreements applying to a data set:
From the combination of purpose with master data origin, we could have several data sharing agreements for the same data set.
Marie-Aude
jennifertemple
·4 years ago · EditedThis is a fantastic thread - who else wants @brad.stone to present at an upcoming data citizens event?
I know I would love to learn more as we are at the beginning stages of designing something similar.
marieaudemagarshack
OP4 years ago · EditedHello Brad, when you write about “pre-curated data sets”, does that include the confidentiality level and PII (Personally Identifiable Information) of the data elements in the data set? I wonder if the Data Sharing Request allow your data stewards to improve the meta-data quality, through adding this piece of information when it is missing. From our side, we are looking at workflows branching to different roles, depending on the confidentiality and PII of the data elements in the data set. Is this a feature included in the workflows you have developed, or are Data Stewards fully in charge of approving Data Sharing Requests?
Best regards
bradstone
·4 years ago · EditedHi Marie,
We have four confidentiality levels (public, internal, confidential, restricted) that are set as an attribute on the Business Term in the glossary. We also have two levels of Privacy (Personal Information, Sensitive Personal Information) that is also an attribute on the Business Term in the glossary.
We don’t require all meta-data to be included on every Business Term entry in the glossary - it is up to the Business Team to maintain their glossary and they decide how much effort to put into the completeness of the data. However, when a new data set is curated because there is an identified need, every field in the data set is matched to the appropriate Business Term. This allows us to:
These activities happen most frequently with the Glossary Editors who are on the Data Steward’s team.
As the data set is requested, it is reviewed by Data Stewards for a specific use (e.g. we want to use this data set for this mobile app). If a glossary entry is not resonating with the Data Steward (e.g. definition is not clear or perhaps the Data Steward disagrees with the confidentiality level, …) then the glossary entry is updated.
Thus, the more the data set is requested, the better the quality of the meta-data. This means that we are spending time and effort on the resources that are used most.
The Privacy attribute is handled slightly differently. Our Chief Privacy Officer (CPO) has reviewed every Business Term in every glossary and marked those that he believes are PI or Sensitive PI. Every time there is a new term added to a glossary, a workflow is automatically kicked off to notify the CPO so he can review the new term. We only started doing the privacy part this year. We plan on training our Data Stewards early next year on how they can use this Privacy information along with their own data knowledge to make the best decisions about data usage.
bradstone
·4 years ago · EditedHi David,
Our work on Data Sharing Agreements predates Collibra Catalog - and even Collibra v5.x, so all of it is done with custom workflows. It also has a custom skin that is fully responsive to allow anyone to make a request from the web or mobile device.
A Data Requester selects from pre-curated data sets that are from sources such as our Data Warehouse (Oracle), Data Virtualization (Dremio), API (various technologies), Single Sign-On (SAML/CAS), and soon Syslog streams. If someone doesn’t find what they need using these technologies, we start a request to curate a new data set.
There also is an option for a Data Requester to request an ala-cart set for a one-time / periodic / or Tableau data extract. They do this by browsing / searching the data glossaries and adding business terms to the cart. From there, we have a Data Concierge meet with them to determine the best data for their needs.
Collibra is a great data governance engine, has rich API functionality, and complete (albeit complex) workflows that have allowed us to create this solution, but it doesn’t come out of the box.
marieaudemagarshack
OP4 years ago · EditedHello, thank you Brad for sharing this feed-back. On our side, we want to go further than a 1st approach, which was similar to the Data Sharing Request approach:
At that time, the implementation of Collibra had just started, we built the process out of this tool.
Now, we want to leverage Collibra’s features to represent the rules that must drive the authorization, the parties involved in authorizing access then in creating the access to the data in the data lake, and the characteristics of the requested data set.
Best regards
bradstone
·4 years ago · EditedAt BYU, we have been using Data Sharing Agreements for over seven years. The process has matured quite a bit over that time. Collibra and associated workflows have allowed us to significantly reduce the time from first request through review through provisioning access.
After doing Data Sharing Agreements for a few years, we discovered that we really needed two views of the request:
A Data Sharing Request (DSR) which records the request from the requesters point of view. It includes the data that they need, their explanation of their use, etc. The requester does not need to know who has stewardship over the data that they request. Collibra will figure that out automatically.
One or more Data Sharing Agreements (DSAs) are generated from the DSR - one for each area of stewardship. These DSAs allow each of the Data Stewards to see the overall request, plus how their data will contribute to the success of the request.
–
In addition to getting the best minds looking at each request, we have found that Data Sharing Requests / Agreements are the secret sauce that keeps all of our meta-data up-to-date. The request process and workflows build /update the relationships between business terms, the technical data models, and the project that needs the data.
When we started doing Data Sharing Agreements in 2014, they were paper / .pdf based. They took about 5 weeks to complete and we completed 18 in the first year.
Today, the process is automated in Collibra. Each request takes about 3 days on average to complete and we are averaging 49 DSAs per month (high of 79 in June 2021). Our goal is to process and provision all requests within one business day, and quite a few of the simpler requests do this today.
Most of our data sets are curated and each data element has a relationship to a Business Term. When someone requests data that is not yet in a data set, we do just-in-time Data Governance and document the data before it is shared.
We haven’t used this process with a data lake per se, but I wonder if just-in-time approach would work in that situation or not. If not, I would be interested in how such data could be documented.
Hope that this helps.
bradstone
·4 years ago · EditedAt BYU, we have been using Data Sharing Agreements for over seven years. The process has matured quite a bit over that time. Collibra and associated workflows have allowed us to significantly reduce the time from first request through review through provisioning access.
After doing Data Sharing Agreements for a few years, we discovered that we really needed two views of the request:
A Data Sharing Request (DSR) which records the request from the requesters point of view. It includes the data that they need, their explanation of their use, etc. The requester does not need to know who has stewardship over the data that they request. Collibra will figure that out automatically.
One or more Data Sharing Agreements (DSAs) are generated from the DSR - one for each area of stewardship. These DSAs allow each of the Data Stewards to see the overall request, plus how their data will contribute to the success of the request.
–
In addition to getting the best minds looking at each request, we have found that Data Sharing Requests / Agreements are the secret sauce that keeps all of our meta-data up-to-date. The request process and workflows build /update the relationships between business terms, the technical data models, and the project that needs the data.
When we started doing Data Sharing Agreements in 2014, they were paper / .pdf based. They took about 5 weeks to complete and we completed 18 in the first year.
Today, the process is automated in Collibra. Each request takes about 3 days on average to complete and we are averaging 49 DSAs per month (high of 79 in June 2021). Our goal is to process and provision all requests within one business day, and quite a few of the simpler requests do this today.
Most of our data sets are curated and each data element has a relationship to a Business Term. When someone requests data that is not yet in a data set, we do just-in-time Data Governance and document the data before it is shared.
We haven’t used this process with a data lake per se, but I wonder if just-in-time approach would work in that situation or not. If not, I would be interested in how such data could be documented.
Hope that this helps.
davidbotzenhart
·4 years ago · EditedHi Brad, thank you for sharing this information. This is very helpful. Question though, are you using the OOTB Data Request feature in Collibra for this? That is only limited to putting the Data Sets in the cart. Are people “Building” a data set in Collibra when they don’t find a existing data set that meets their needs? I’m new to Collibra if you can’t tell.
sergiuszrotman
·4 years ago · EditedHi Marie-Aude, if possible, I would be happy to discuss DSAs, as well as broader initiatives we take now for managing Access Policies in Collibra. I have already contacted our Customer Advisory team.
marieaudemagarshack
OP4 years ago · EditedHello Sergiusz, thank you for your reply. I will be very happy to discuss further the DSA topic with you.
Best regards
sergiuszrotman
·4 years ago · EditedHi Marie-Aude, if possible, I would be happy to discuss DSAs, as well as broader initiatives we take now for managing Access Policies in Collibra. I have already contacted our Customer Advisory team.
Former User
·4 years ago · EditedIt would best serve the purpose of the community if you contributions could be shared in this thread.
marieaudemagarshack
OP4 years ago · EditedHello Robert, thank you very much for sharing your approach! I keep it in mind. If the ‘Data Sharing Agreement’ asset type is not adequate for displaying the evidence that a data set can be shared, we will probable build a specific asset type.
Best regards
Former User
·4 years ago · EditedHi Marie-Aude, … we drive a concept that we name “Data Sharing Terms”. Our main assumption is, that no data comes without any context, that describes how data can be used. Context can be legal , contractual as well as compliance topics and regulations. (GDPR, consent agreements, country regulartion, etc. ) … The concept that we want to enalbe is data democratization through informed use of data. We therefore implemted the concept of “Data Sharing Terms” bound to a data usage within Collibra. Regards Robert