Registering Data Source, Catalog, Anonymize Data, Auto Classification
Hello,
My team is in the process of integrating our SQL environment with Collibra (using the SQL JDBC driver), and we have some questions. Thought this would be a great forum to get input from other data citizens who have implemented this or who are looking to implement this soon.
-
If you decide to import sample data via an API, can you still do auto classification or you will have to trigger the classification manually? I would like to understand the limitations/benefits with using this over selecting the ‘store sample data’ option when registering a data source.
-
If there is a concern with storing data on the cloud, I found that you can enable the ‘anonymize data’ option. However, this will disable auto data classification. What is the recommended solution if we want to use anonymize data option but still classify our data for building other capabilities such as knowledge graphs?
-
There are talks on Collibra about introducing ‘Edge’ where sample data might not be stored on the cloud. When is this feature being released and how different is it from the ‘anonymize data’ option?
Thanks in advance,
Akira
arvindsingh
·5 years ago · EditedLimitation: It might take more time to ingest a large database, where API might get it done in hours.
Benefit: No maintenance/support which you might have to do on API integration.
Currently, if you enable the data anonymization process you can no longer use Automatic Data Classification. May be use APIs to assign classes manually without using anonymized sample data. Personally, I can’t see a reason to store anonymized data if you are not going to use it for classification.
This was supposed to come in January, 2021 and now, it has been pushed back to April, 2021. Most of upcoming features of Collibra are mystery until released. Impossible to find any documentation on these.