Abstract: Disclosed subject matter is related to digital trust including method and apparatus for determining a digital identity trust index in digital ecosystems. The method of a computing device includes retrieving, contextual data related to a subscriber from data repositories upon receiving a data request and generating intermediate analytics results by performing federated analytics on the contextual data. Sensitive contextual data is secured using a noise perturbed feature based technique, before performing the federated analytics. Further, a trust index determining server receives intermediate analytics results from the computing device and generates a prediction result using a machine learning model trained to predict using an optimal feature set. The prediction result thus generated is considered as the digital identity trust index of the subscriber. The present disclosure determines the digital identity trust index for wide variety of applications, accurately and in real-time, while maintaining user privacy and data security. FIGs. 2A and 2D
1. A method of generating analytics results for determining a digital identity trust index in digital ecosystem, wherein the digital ecosystem comprises a plurality of computing devices (103) associated with a corresponding plurality of data repositories (101), the method comprising: receiving, by a computing device (103), a data request from a trust index determining server (107), wherein the data request is related to a subscriber; retrieving, by the computing device (103), contextual data (207) related to the subscriber, from a corresponding data repository (101) of the plurality of data repositories (101), wherein the contextual data (207) comprises at least one of sensitive contextual data and non-sensitive contextual data; generating, by the computing device (103), secure sensitive contextual data (209) by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique; and generating, by the computing device (103), one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data (209) and the non-sensitive contextual data, using a third party analytics agent (105) associated with the computing device (103), wherein the one or more intermediate analytics results are transmitted to the trust index determining server (107), for determining a digital identity trust index in digital ecosystems.
2. The method as claimed in claim 1, wherein the degree of sensitivity of each contextual element is determined by: determining a predefined category of each of the contextual element, wherein each predefined category is associated with a degree of sensitivity; and assigning each contextual element with the degree of sensitivity of the corresponding predefined category.
3. The method as claimed in claim 1, wherein the degree of sensitivity of each contextual element is quantified into one of “Critical”, “High”, “Medium” and “Low”.
4. The method as claimed in claim 1, wherein the noise perturbed feature based technique comprises: dividing at least one of the sensitive contextual data and the non-sensitive contextual data into plurality of matrices, based on total number of the contextual elements of at least one of the sensitive contextual data and the non-sensitive contextual data, wherein the contextual elements of the sensitive contextual data are referred as sensitive contextual elements; clustering a plurality of sensitive contextual elements among total number of sensitive contextual elements, into a plurality of clusters, based on a predefined clustering technique, wherein a centroid value is computed for each of the plurality of clusters; and replacing a value of, each of the plurality of sensitive contextual elements with a combined value obtained by combining the centroid value of each of the corresponding plurality of clusters and a first corresponding random noise, wherein the first random noise is selected for each of the plurality of sensitive contextual elements from a Laplace distribution, based on a degree of sensitivity of each of the plurality of sensitive contextual elements; and each of rest of the plurality of sensitive contextual elements among the total number of sensitive contextual elements with a combined value obtained by combining original value of each corresponding sensitive contextual element and a second corresponding random noise, wherein the second random noise is selected for each of rest of the plurality of sensitive contextual elements from the Laplace distribution, based on a degree of sensitivity of each of the rest of the plurality of sensitive contextual elements.
5. The method as claimed in claim 1, wherein the subscriber is one of a customer, a service provider or a small/medium enterprise.
6. The method as claimed in claim 1, wherein generating the one or more intermediate analytics results by performing federated analytics comprises: identifying, one or more candidate features required for determining the digital identity trust index of the subscriber, and one or more sub-features related to each of the one or more candidate features, from at least one of the secure sensitive contextual data (209) and the non-sensitive contextual data, based on an optimal feature set; and determining a normalized value with respect to a predefined scale, for each of the one or more candidate features using the one or more sub-features related to each of the one or more candidate features, wherein the one or more candidate features and the normalized value corresponding to each of the one or more candidate features constitute the one or more intermediate results.
7. A method of determining a digital identity trust index in digital ecosystems, the method comprising: receiving, by a trust index determining server (107), one or more intermediate analytics results from a plurality of computing devices (103) associated with a corresponding plurality of data repositories (101), wherein the one or more intermediate results are related to a subscriber whose digital identity trust index is to be determined, wherein the one or more intermediate results comprises one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features; selecting, by the trust index determining server (107), a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set; and generating, by the trust index determining server (107), a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber, wherein the machine learning model is trained to generate the prediction result based on the optimal feature set.
8. The method as claimed in claim 7, wherein the optimal feature set, the key feature of the one or more candidate features, and the rest of the one or more candidate features are determined during a training phase of the machine learning model, based on historical data related to an application for which the digital identity trust index of the subscriber is being determined.
9. The method as claimed in claim 7, wherein the received one or more intermediate analytics results are generated by performing federated analytics on at least one of secure sensitive contextual data (209) and non-sensitive contextual data, related to the subscriber, wherein the secure sensitive contextual data (209) is obtained by applying a noise perturbed feature based technique on sensitive contextual data related to the subscriber.
10. The method as claimed in claim 7 further comprises determining, by the trust index determining server (107), an updated digital identity trust index, periodically, by determining an exponential moving average of currently determined digital identity trust index based on updated contextual data, and the previously determined digital identity trust index.
11. The method as claimed in claim 7, wherein the subscriber is one of a customer, a service provider or a small/medium enterprise.
12. A computing device (103) for generating analytics results for determining a digital identity trust index in digital ecosystem, wherein the digital ecosystem comprises a plurality of computing devices (103) associated with a corresponding plurality of data repositories (101), the computing device (103) comprising: a processor (109); and a memory (113) communicatively coupled to the processor (109), wherein the memory (113) stores the processor (109)-executable instructions, which, on execution, causes the processor (109) to: receive a data request from a trust index determining server, wherein the data request is related to a subscriber; retrieve contextual data (207) related to the subscriber, from a corresponding data repository of the plurality of data repositories (101), wherein the contextual data (207) comprises at least one of sensitive contextual data and non-sensitive contextual data; generate secure sensitive contextual data (209) by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique; and generate one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data (209) and the non-sensitive contextual data, using a third party analytics agent (105) associated with the computing device (103), wherein the one or more intermediate analytics results are transmitted to the trust index determining server, for determining a digital identity trust index in digital ecosystems.
13. The computing device (103) as claimed in claim 12, wherein the processor (109) determines degree of sensitivity of each contextual element by: determining a predefined category of each of the contextual element, wherein each predefined category is associated with a degree of sensitivity; and assigning each contextual element with the degree of sensitivity of the corresponding predefined category.
14. The computing device (103) as claimed in claim 12, wherein the degree of sensitivity of each contextual element is quantified into one of “Critical”, “High”, “Medium” and “Low”.
15. The computing device (103) as claimed in claim 12, wherein to perform the noise perturbed feature based technique, the processor (109) is configured to: divide at least one of the sensitive contextual data and the non-sensitive contextual data into plurality of matrices, based on total number of the contextual elements of at least one of the sensitive contextual data and the non-sensitive contextual data, wherein the contextual elements of the sensitive contextual data are referred as sensitive contextual elements; cluster a plurality of sensitive contextual elements among total number of sensitive contextual elements, into a plurality of clusters, based on a predefined clustering technique, wherein a centroid value is computed for each of the plurality of clusters; and replace a value of, each of the plurality of sensitive contextual elements with a combined value obtained by combining the centroid value of each of the corresponding plurality of clusters and a first corresponding random noise, wherein the first random noise is selected for each of the plurality of sensitive contextual elements from a Laplace distribution, based on a degree of sensitivity of each of the plurality of sensitive contextual elements; and each of rest of the plurality of sensitive contextual elements among the total number of sensitive contextual elements with a combined value obtained by combining original value of each corresponding sensitive contextual element and a second corresponding random noise, wherein the second random noise is selected for each of rest of the plurality of sensitive contextual elements from the Laplace distribution, based on a degree of sensitivity of each of the rest of the plurality of sensitive contextual elements.
16. The computing device (103) as claimed in claim 12, wherein the subscriber is one of a customer, a service provider or a small/medium enterprise.
17. The computing device (103) as claimed in claim 12, wherein to generate the one or more intermediate analytics results by performing federated analytics, the processor (109) is configured to: identify one or more candidate features required for determining the digital identity trust index of the subscriber, and one or more sub-features related to each of the one or more candidate features, from at least one of the secure sensitive contextual data (209) and the non-sensitive contextual data, based on an optimal feature set; and determine a normalized value with respect to a predefined scale, for each of the one or more candidate features using the one or more sub-features related to each of the one or more candidate features, wherein the one or more candidate features and the normalized value corresponding to each of the one or more candidate features constitute the one or more intermediate results.
18. A trust index determining server (107) for determining a digital identity trust index in digital ecosystems, the trust index determining server (107) comprising: a processor (115); and a memory (119) communicatively coupled to the processor (115), wherein the memory (119) stores the processor-executable instructions, which, on execution, causes the processor (115) to: receive one or more intermediate analytics results from a plurality of computing devices (103) associated with a corresponding plurality of data repositories (101), wherein the one or more intermediate results are related to a subscriber, whose digital identity trust index is to be determined, wherein the one or more intermediate results comprises one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features; select a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set; and generate a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber, wherein the machine learning model is trained to generate the prediction result based on the optimal feature set.
19. The trust index determining server (107) as claimed in claim 18, wherein the optimal feature set, the key feature of the one or more candidate features, and the rest of the one or more candidate features are determined during a training phase of the machine learning model, based on historical data related to an application for which the digital identity trust index of the subscriber is being determined.
20. The trust index determining server (107) as claimed in claim 18, wherein the received one or more intermediate analytics results are generated by performing federated analytics on at least one of secure sensitive contextual data (209) and non-sensitive contextual data, related to the subscriber, wherein the secure sensitive contextual data (209) is obtained by applying a noise perturbed feature based technique on sensitive contextual data related to the subscriber.
21. The trust index determining server (107) as claimed in claim 18, wherein the processor (115) is further configured to determine an updated digital identity trust index, periodically, by determining an exponential moving average of currently determined digital identity trust index based on updated contextual data, and the previously determined digital identity trust index.
22. The trust index determining server (107) as claimed in claim 18, wherein the subscriber is one of a customer, a service provider or a small/medium enterprise. , Description:TECHNICAL FIELD The present subject matter is related, in general to digital trust and more particularly, but not exclusively to method and system for determining digital identity trust index in digital ecosystems. BACKGROUND In today’s digital world, rate of economic growth depends on emergence of new business opportunities and level of trust among participating entities. Generation of huge amount of digital data has led to creation of diverse public and private data repositories and platforms that support building business applications providing rich insights by data analytics. However, for co-creation and sustained usage of new data driven applications, and for digital transformation of economies, there exists a need to build trust among the participating entities, such as customers, service providers and Small/Medium Enterprises (SMEs). Existing verification processes are built on submission of physical documents, which have a limited scope with respect to data sources. Moreover, verification based on submitted physical documents is time-consuming, unreliable and based on human interpretation of data records, which leads to high number of fraud cases, poor services and reduced digital trust. Currently, for ensuring digital trust, scores such as Credit Information Bureau (India) Limited (CIBIL®) and Zhima® credit associate a number to a customer, which defines level of trust and credit worthiness of the customer. Banks and other organizations provide incentives to the customer based on the score. However, the CIBIL® and Zhima® credit scores are limited only to financial industry, due to which, CIBIL® score relies only on financial transaction history of the customer as source data, and Zhima® credit score relies on financial transaction history as well as behavioural analysis of the customers, while using specific online websites, as the source data. Apart from CIBIL® and Zhima® credit scores mentioned above, currently there are other techniques in use for computing identity and trust worthiness of online users. Such techniques integrate user data retrieved from social media, e-commerce and other public repositories along with subscriber data provided directly by the user. Upon retrieving the user data, these techniques compare correctness of subscriber data values with the user data retrieved from limited number of public repositories, to determine trust score for the user. However, such techniques are completely dependent on correctness of the subscriber data provided by the user rather than contextual interpretation of user data retrieved from public repositories, which leads to inaccuracy in trust score determination. Moreover, such techniques retrieve user data from limited number of public repositories, which may lead to the risk of missing out on important information related to the user. Yet another existing technique allows a user to select a service provider from a group, based on trust scores associated with each of the service providers. In this existing technique, the service provider trust score is determined based on user reputation score and a general reputation score. User reputation score is a score computed based on current user’s perception of the service provider, and the general reputation score is a score computed based on generic perception of other users who have already received services of the service provider. Therefore, this technique completely relies on the perception of users to determine the service provider trust score, which is just one dimension for determining the trust score. Such single dimensional computations have an extremely narrow scope, and are inaccurate due to their sole dependency on perception of users, rather than user data and facts available in public repositories. Therefore, there exists a need to accurately determine the digital identity trust index, for wide variety of applications in digital ecosystems. The information disclosed in this background of the disclosure section is only for enhancement of understanding of the general background of the invention and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art. SUMMARY Disclosed herein is a method of generating analytics results for determining a digital identity trust index in digital ecosystem. The digital ecosystem comprises a plurality of computing devices associated with a corresponding plurality of data repositories. The method includes receiving, by a computing device, a data request from a trust index determining server, wherein the data request is related to a subscriber. Further, the method includes retrieving contextual data related to the subscriber, from a corresponding data repository of the plurality of data repositories. The contextual data comprises at least one of sensitive contextual data and non-sensitive contextual data. Subsequently, the method includes generating secure sensitive contextual data by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique. Upon generating the secure sensitive contextual data, the method includes generating one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data and the non-sensitive contextual data, using a third party analytics agent associated with the computing device. The one or more intermediate analytics results are transmitted to the trust index determining server, for determining a digital identity trust index in digital ecosystems. Further, the present disclosure discloses a method of determining a digital identity trust index in digital ecosystems. The method includes receiving, by a trust index determining server, one or more intermediate analytics results from a plurality of computing devices associated with a corresponding plurality of data repositories. The one or more intermediate results are related to a subscriber whose digital identity trust index is to be determined. The one or more intermediate results include one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. Further, the method includes selecting a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set. Finally, the method includes generating a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber. The machine learning model is trained to generate the prediction result based on the optimal feature set. Further, the present disclosure includes a computing device for generating analytics results for determining a digital identity trust index in digital ecosystem. The digital ecosystem includes a plurality of computing devices associated with a corresponding plurality of data repositories. The computing device includes a processor and a memory communicatively coupled to the processor. The memory stores the processor-executable instructions, which, on execution, causes the processor to receive a data request from a trust index determining server, wherein the data request is related to a subscriber. Further, the processor retrieves contextual data related to the subscriber, from a corresponding data repository of the plurality of data repositories. The contextual data comprises at least one of sensitive contextual data and non-sensitive contextual data. Further, the processor generates secure sensitive contextual data by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique. Upon generating the secure sensitive contextual data, the processor generates one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data and the non-sensitive contextual data, using a third party analytics agent associated with the computing device. The one or more intermediate analytics results are transmitted to the trust index determining server, for determining a digital identity trust index in digital ecosystems. Furthermore, the present disclosure discloses a trust index determining server for determining a digital identity trust index in digital ecosystems. The trust index determining server comprises a processor and a memory communicatively coupled to the processor. The memory stores the processor-executable instructions, which, on execution, causes the processor to receive one or more intermediate analytics results from a plurality of computing devices associated with a corresponding plurality of data repositories. The one or more intermediate results are related to a subscriber, whose digital identity trust index is to be determined. The one or more intermediate results comprises one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. Thereafter, the processor selects a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set. Finally, the processor generates a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber. The machine learning model is trained to generate the prediction result based on the optimal feature set. The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE ACCOMPANYING DIAGRAMS The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The same numbers are used throughout the figures to reference like features and components. Some embodiments of system and/or methods in accordance with embodiments of the present subject matter are now described, by way of example only, and with reference to the accompanying figures, in which: FIG.1A shows an exemplary architecture for generating analytics results and determining a digital identity trust index in a digital ecosystem in accordance with some embodiments of the present disclosure; FIG.1B shows a brief block diagram of a computing device in accordance with some embodiments of the present disclosure; FIG.1C shows a brief block diagram of a trust index determining server in accordance with some embodiments of the present disclosure; FIG.2A shows a detailed block diagram of a computing device for generating analytics results for determining a digital identity trust index in digital ecosystem in accordance with some embodiments of the present disclosure; FIG.2B shows an exemplary matrix comprising sensitive and non-sensitive data in accordance with some embodiments of the present disclosure; FIG.2C shows an exemplary matrix with clusters in accordance with some embodiments of the present disclosure; FIG.2D shows a detailed block diagram of a trust index determining server for determining a digital identity trust index in digital ecosystems in accordance with some embodiments of the present disclosure; FIG.2E and FIG.2F show exemplary scenarios for determining a digital identity trust index in a digital ecosystem in accordance with some embodiments of the present disclosure; FIG.3A shows a flowchart illustrating a method of generating analytics results for determining a digital identity trust index in digital ecosystem in accordance with some embodiments of the present disclosure; FIG.3B shows a flowchart illustrating a method of determining a digital identity trust index in digital ecosystems in accordance with some embodiments of the present disclosure; and FIG.4 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure. It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown. DETAILED DESCRIPTION In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the scope of the disclosure. The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises… a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or method. Disclosed herein are method and apparatus for determining digital identity trust index in digital ecosystems for a wide variety of applications. The method includes receiving, by a computing device, a data request from a trust index determining server. In some embodiments, the data request may be related to a subscriber. In some embodiments, the subscriber may be of a customer, a service provider or a small/medium enterprise. Further, the computing device may retrieve contextual data related to the subscriber, from a corresponding data repository of the plurality of data repositories. As an example, the contextual data may be financial data, tax related data, identity verification data, police and judicial data, social media data, health related data and the like, related to the subscriber. The contextual data may include at least one of sensitive contextual data and non-sensitive contextual data. In the present disclosure, the computing devices retrieve the contextual data from the corresponding data repositories and provide the contextual data to the corresponding third party analytics agents. In some embodiments, in order to secure privacy of the sensitive contextual data related to the subscriber, the present disclosure employs a noise perturbed feature based technique to add noise to the sensitive contextual data, prior to providing the sensitive contextual data to the corresponding third party analytics agents. The third party analytics agents may perform federated analytics on the non-sensitive contextual data and secure sensitive contextual data, which results in one or more intermediate analytics results. The one or more intermediate results may include, but not limited to, one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. In some embodiments, the plurality of computing devices may transmit the one or more intermediate results to a trust index determining server for further processing. The trust index determining server may generate a prediction result based on each of the one or more intermediate results, for determining the digital identity trust index of the subscriber. In some embodiments, the prediction result may be generated using a machine learning model which is trained to generate the prediction result based on an optimal feature set. The present disclosure allows determination of digital trust index for wide variety of applications such as merchant trust score, data trust score for information sharing, customer trust score for financial services, logistics service provider trust score, demographic trust score for smart city planning, e-commerce service provider trust score, employee onboarding trust score, mobility service provider trust score and the like. Further, since the present disclosure determines the digital identity trust index by performing contextual interpretation of the user data retrieved from wide variety of public and private data repositories, the scope of verification is broad, which in turn results in accurate comprehension of trust levels of the customer, service provider or the small/medium enterprises. Such accurate comprehension of trust levels reduces the number of fraud cases and provides early risk management to businesses. Further, the noise perturbed feature based technique used in the present disclosure anonymizes the sensitive contextual data to an extent that it ensures confidentiality of the sensitive contextual data, but also allows to perform meaningful data analytics, by the third party analytics agents. Moreover, since the present disclosure deploys the third party analytics agents at the respective location of each of the public and private data repositories, contextual data of the user remains in its locked-down location, and intermediate analytics results are generated locally. This eliminates the need to transmit large amount of the contextual data over the communication network, thereby preventing exposure of the contextual data to security breaches and data theft. Moreover, transmitting only the intermediate analytics results to the trust index determining server eliminates over utilization of the available bandwidth, thereby enabling bandwidth conservation. Further, since the present disclosure generates prediction result based on optimal feature set, the present disclosure reduces processing complexity and enhances the speed of determination of the digital trust identity index, which helps in performing/providing services dependent on digital identity trust index of the subscriber, in real-time.. Moreover, the machine learning model used in the present disclosure automates the process of determining digital identity trust index, which means human interpretation of records is eliminated, thereby improving accuracy of the digital identity trust index. Overall, the present disclosure determines the digital identity trust index for wide variety of applications, accurately and in real-time, while maintaining user privacy and data security. In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense. FIG.1A shows an exemplary architecture for generating analytics results and determining a digital identity trust index in a digital ecosystem in accordance with some embodiments of the present disclosure. The architecture 100 comprises a data repository 1 1011 to a data repository n 101n (collectively referred to as plurality of data repositories 101), a computing device 1 1031 to a computing device n 103n (collectively referred to as plurality of computing devices 103), a third party analytics agent 1 1051 to a third party analytics agent n 105n (collectively referred to as plurality of third party analytics agents 105), and a trust index determining server 107. The plurality of data repositories 101 may be public data repositories or private repositories. As an example, the plurality of data repositories 101 may include, but not limited to, financial transactions repository, Tax/Goods and Services Tax (GST) records repository, insurance data repository, investment records repository, social media data repository, demographic statistics repository, identity verification repository, health stack information repository, police records repository, Central Bureau of Investigation (CBI) records repository, judiciary records repository, online customer feedback repository, service quality records repository, blockchain records repository and telecom data repository. In some embodiments, each of the plurality of data repositories 101 may be associated with each of the corresponding plurality of computing devices 103 as shown in the FIG.1A. In some other embodiments, the plurality of data repositories 101 may be configured within the corresponding plurality of computing devices 103. As an example, the plurality of computing devices 103 may include, but not limited to, a laptop, a computer, a desktop, a mobile phone, a tablet phone and the like. Further, each of the plurality of computing devices 103 may be associated with the corresponding plurality of third party analytics agents 105 as shown in the FIG.1A. In some embodiments, the plurality of third party analytics agents 105 may be configured within the corresponding plurality of computing devices 103. Each of the plurality of third party analytics agents 105 are locally configured i.e. configured at the same location as that of the corresponding plurality of computing devices 103. In some embodiments, each of the plurality of computing devices 103 may be associated with the trust index determining server 107 via a communication network (not shown in the FIG.1A). The communication network may be a wireless communication network. In some embodiments, the trust index determining server 107 may be configured remotely. In some other embodiments, the trust index determining server 107 may be configured in a cloud environment. Each of the plurality of computing devices 103 comprises a processor 109, an Input/output (I/O) interface 111 and a memory 113 as shown in the FIG.1B. For the ease of understanding, the present disclosure is explained further for a single computing device. The same method is followed by each of the plurality of computing devices 103. In an embodiment, the I/O interface 111 of the computing device 103 may receive a data request related to a subscriber, from the trust index determining server 107. In some embodiments, the subscriber may include, but not limited to, a customer, a service provider or a Small/Medium Enterprise (SME). Further, the processor 109 of the computing device 103 may retrieve contextual data related to the subscriber from the corresponding data repository 101. In some embodiments, the contextual data may include, but not limited to, sensitive contextual data and non-sensitive contextual data. In some embodiments, the contextual data may be retrieved upon receiving a consent from the subscriber to access data records related to the subscriber. In some embodiments, the consent of the subscriber may be implicit in nature, in order to avail a digital identity trust index. Further, the processor 109 of the computing device 103 may generate secure sensitive contextual data using a noise perturbed feature based technique on the retrieved sensitive contextual data. Thereafter, the processor 109 of the computing device 103 may provide the secure sensitive contextual data and the non-sensitive contextual data to the corresponding third party analytics agent 105, to generate one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data and the non-sensitive contextual data. The processor 109 of the computing device 103 may then transmit the one or more intermediate analytics results to the trust index determining server 107 for determining a digital identity trust index. The trust index determining server 107 comprises a processor 115, an Input/output (I/O) interface 117 and a memory 119 as shown in FIG.1C. In an embodiment, the I/O interface 117 of the trust index determining server 107 may receive one or more intermediate analytics results from the plurality of computing devices 103 associated with the corresponding plurality of data repositories 101. In some embodiments, the one or more intermediate results may include, but not limited to, one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. The processor 115 of the trust index determining server 107 may determine the digital identity trust index of the subscriber by providing the one or more intermediate results to a machine learning model of the trust index determining server 107. The digital identity trust index thus determined may be stored in the memory 119 of the trust index determining server 107. FIG.2A shows a detailed block diagram of a computing device for generating analytics results for determining a digital identity trust index in digital ecosystem in accordance with some embodiments of the present disclosure. In some implementations, the computing device 103 may include data 203 and modules 205. As an example, the data 203 of the computing device 103 are stored in the memory 113 configured in the computing device 103 as shown in the FIG.2A. In one embodiment, the data 203 of the computing device 103 may include contextual data 207, secure sensitive contextual data 209, intermediate result data 211 and other data 213. In the illustrated FIG.2A, modules 205 of the computing device 103 are described herein in detail. In some embodiments, the data 203 of the computing device 103 may be stored in the memory 113 of the computing device 103 in the form of various data structures. Additionally, the data 203 of the computing device 103 can be organized using data models, such as relational or hierarchical data models. In some embodiments, the other data 213 may store data, including temporary data and temporary files, generated by the modules 205 of the computing device 103 for performing the various functions of the computing device 103. In some embodiments, the data 203 of the computing device 103 stored in the memory 113 may be processed by the modules 205 of the computing device 103. The modules 205 may be stored within the memory 113 of the computing device 103. In an example, the modules 205 communicatively coupled to the processor 109 configured in the computing device 103, may also be present outside the memory 113 as shown in FIG.2A and implemented as hardware. As used herein, the term modules refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or group) and memory that execute one or more software or firmware programs, a combinational logic circuit, and/or other suitable components that provide the described functionality. In an embodiment, the modules 205 of the computing device 103 may include, for example, a receiving module 221, data retrieving module 223, a secure data generating module 225, an intermediate result generating module 227, a transmitting module 229 and other modules 231. The other modules 231 may be used to perform various miscellaneous functionalities of the computing device 103. It will be appreciated that such aforementioned modules 205 of the computing device 103 may be represented as a single module or a combination of different modules. In some embodiments, the receiving module 221 may receive a data request from a trust index determining server 107 associated with the computing device 103. The data request may be related to a subscriber. In some embodiments, the receiving module 221 may receive the data request from the trust index determining server 107, when the subscriber wishes to avail a service related to a business application. In some embodiments, the trust index determining server 107 may send the data request to some of the plurality of computing devices 103 i.e. to some of the plurality of data repositories based on a kind of service which the subscriber wishes to avail. In some other embodiments, the trust index determining server 107 may send the data request to each of the plurality of computing devices 103, irrespective of the kind of service that the subscriber wishes to avail. In yet other embodiments, the trust index determining server 107 may send the data request to some of the plurality of computing devices 103 i.e. to some of the plurality of data repositories based on the kind of subscriber. For example, when the subscriber is a customer, the data request may be sent to tax records repository, as the tax records are relevant for the customer. However, when the subscriber is a service provider or a Small/Medium Entity (SME), the data request may be sent to GST records repository since the GST records are relevant for the service providers and the SMEs, and not the tax records. In some embodiments, the data retrieving module 223 may retrieve the contextual data 207 related to the subscriber, from a corresponding data repository 101 of the plurality of data repositories 101. In some embodiments, the contextual data may include, but not limited to, at least one of sensitive contextual data and non-sensitive contextual data. Sensitive contextual data may be the contextual data 207 related to the subscriber, which is confidential in nature. As an example, in the context of financial transactions, data related to date and time of financial transactions, transaction amount, and Automated Teller Machine (ATM)/Point Of Sale (POS)/bank location of the transaction and the like, may be regarded as sensitive contextual data of the subscriber. Non-sensitive contextual data may be the contextual data 207 related to the subscriber, which is non-confidential in nature, and easily available in public. As an example, name, age, gender, educational qualification and the like, may be regarded as non-confidential data of the subscriber. In some embodiments, the secure data generating module 225 may generate secure sensitive contextual data 209 by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique. As an example, if the sensitive contextual data is related to financial transactions, the contextual elements of the sensitive contextual data may be “date and time of the transactions”, “transaction amount”, and “location of the transaction” and the like. As an example, if the sensitive contextual data is related to health records, the contextual elements of the sensitive contextual data may be “hospital location”, “disease diagnosed”, and “date and time of doctor visit” and the like. In some embodiments, to perform the noise perturbed feature based technique based on the degree of sensitivity of each contextual element, the secure data generating module 225 may divide at least one of the sensitive contextual data and the non-sensitive contextual data into plurality of matrices. In some embodiments, the secure data generating module 225 may divide based on total number of the contextual elements of at least one of the sensitive contextual data and the non-sensitive contextual data. In some embodiments, contextual elements of the sensitive contextual data may be referred as sensitive contextual elements. Further, the secure data generating module 225 may cluster a plurality of sensitive contextual elements among total number of sensitive contextual elements, into a plurality of clusters, based on a predefined clustering technique. As an example, the predefined clustering technique may include, but not limited to, K-means clustering technique. Subsequently, the secure data generating module 225 may compute a centroid value for each of the plurality of clusters. Further, the secure data generating module 225 may replace a value of, each of the plurality of sensitive contextual elements with a combined value obtained by combining the centroid value of each of the corresponding plurality of clusters and a first corresponding random noise. In some embodiments, the secure data generating module 225 may select the first random noise for each of the plurality of sensitive contextual elements from a Laplace distribution, based on the degree of sensitivity of each of the plurality of sensitive contextual elements. In some embodiments, the degree of sensitivity of each contextual element may be quantified into one of “Critical”, “High”, “Medium” and “Low” (also referred as categories of degree of sensitivity). In some embodiments, “Critical” may be the highest order of degree of sensitivity among four categories of the degree of sensitivity mentioned above, followed by the order “High”, “Medium” and “Low”. “Low” is the lowest order of degree of sensitivity among the four categories of the degree of sensitivity mentioned above. In some embodiments, the scale of noise selected for a sensitive contextual element from the Laplace distribution is directly proportional to the degree of sensitivity of the sensitive contextual element. For example, if the degree of sensitivity of the sensitive contextual element is “Critical”, then noise of highest scale is selected from the Laplace distribution. Similarly, if the degree of sensitivity of the sensitive contextual element is “Low”, then noise of lowest scale is selected from the Laplace distribution. In some embodiments, the secure data generating module 225 may determine the degree of sensitivity of each contextual element by initially determining a predefined category of each of the contextual element. Each predefined category may be associated with a degree of sensitivity. Thereafter, the secure data generating module 225 may assign each contextual element of the sensitive contextual data, with the degree of sensitivity of the corresponding predefined category. Further, the secure data generating module 225 may replace a value of each of rest of the plurality of sensitive contextual elements among the total number of sensitive contextual elements with a combined value obtained by combining original value of each corresponding sensitive contextual element and a second corresponding random noise. In some embodiments, the secure data generating module 225 may select the second random noise for each of rest of the plurality of sensitive contextual elements from the Laplace distribution, based on the degree of sensitivity of each of the rest of the plurality of sensitive contextual elements. The data thus generated may be stored as secure sensitive contextual data 209. As an example, consider the following sensitive contextual elements: SC: Sensitive-Critical SH: Sensitive-High SM: Sensitive-Medium SL: Sensitive-Low Further, consider non-sensitive contextual elements that are represented as NSF. Consider, total number of sensitive contextual elements: N=4 and total number of non-sensitive contextual elements: M=3 In some embodiments, the secure data generating module 225 divides the sensitive contextual data and the non-sensitive contextual data into a plurality of matrices, wherein the secure data generating module 225 may determine size of each matrix using the below Equation 1. (N + M) x (N + M)2 ------------- Equation 1 In the above Equation 1, N represents number of sensitive contextual elements in the sensitive contextual data; M represents number of non-sensitive contextual elements in the non-sensitive contextual data. In the above example, N=4 and M=3. Therefore, using the Equation 1, the secure data generating module 225 divides the sensitive contextual data and the non-sensitive contextual data into plurality of matrices, each matrix of size: 7 (Columns) x 49 (Rows) The exemplary matrix with 7 columns and 49 rows is as shown in FIG.2B. In the exemplary matrix shown above, each row denotes a user record and each column denotes the sensitive contextual element/non-sensitive contextual element. Further, the secure data generating module 225 may select plurality of sensitive contextual elements using the below Equation 2. K = floor(sqrt(N)) -------------- Equation 2 In the above Equation 2, N denotes the total number of sensitive contextual elements; and K denotes number of columns of the matrix comprising sensitive contextual elements to be selected. In the above example, since total number of sensitive contextual elements = 4, using the Equation 2, K i.e. number of columns of sensitive contextual elements selected would be 2. Thereafter, the secure data generating module 225 may randomly select K number of columns comprising sensitive contextual elements from the total number of columns comprising the sensitive contextual elements in the matrix. In the above example, consider, the secure data generating module 225 selects columns 1 and 3 of the matrix. Upon using a predefined clustering technique, consider the secure data generating module 225 clusters the values of the sensitive contextual elements present in the selected columns 1 and 3 of the matrix into plurality of clusters, and computes centroid value for each of the plurality of clusters, as shown below: - Column 1: Based on values of sensitive contextual elements: Divided into X = 4 Clusters Centroids of each Cluster - C1,1, C1,2, C1,3, C1,4. - Column 3: Based on values of sensitive contextual elements: Divided into X = 4 Clusters Centroids of each Cluster - C3,1, C3,2, C3,3, C3,4. Size of Cluster: No. of Rows in Cluster: Consider, Cluster 1 = 12, Cluster 2 = 12, Cluster 3 = 12, Cluster 4 = 13. The exemplary four clusters formed in the above example, are as shown in the FIG.2C. Thereafter, the secure data generating module 225 may replace value of each of the plurality of sensitive contextual elements with a combined value obtained by combining the centroid value of each of the corresponding plurality of clusters and a first corresponding random noise. As an example, consider, original value in Column 1: Row 1 is V1,1 as shown in the FIG.2B. The original value V1,1 may be thus replaced by the combined value (C1,1 + N1,1), in which C1,1 denotes centroid value of the corresponding cluster and N1,1 is the first corresponding random noise selected from a Laplace distribution. The first random noise may be selected from the Laplace distribution based on the degree of sensitivity. For instance, if the degree of sensitivity of the sensitive contextual element is “Critical”, then noise of highest scale is selected from the Laplace distribution. Similarly, the secure data generating module 225 replaces each value of the sensitive contextual elements in the columns 1 and 3 of the matrix. The exemplary values that are replaced in columns 1 and 3 of the matrix are as shown in the FIG.2C. Further, for the remaining columns comprising sensitive contextual elements in the matrix, apart from columns 1 and 3, the secure data generating module 225 may replace each value of the sensitive contextual elements of these remaining columns of the matrix, with a combined value obtained by combining original value of each corresponding sensitive contextual element and a second random noise selected from the Laplace distribution. As an example, the remaining columns comprising sensitive contextual elements in the matrix may be columns 2 and 4. As an example, the original value of the column 2: Row 2 is V2,2, as shown in the FIG.2B. The original value V2,2 may be thus replaced by the value (V2,2 + N2,2), which means that second random noise N2,2 is added to the original value V2,2. Similarly, the secure data generating module 225 replaces each value of the sensitive contextual elements in the remaining columns 2 and 4 of the matrix. The exemplary values that are replaced in columns 2 and 4 of the matrix are as shown in the FIG.2C. Further, the intermediate result generating module 227 may generate one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data 209 and the non-sensitive contextual data. The intermediate result generating module 227 may generate the federated analytics using a third party analytics agent 105 associated with the computing device 103. In some embodiments, the intermediate result generating module 227 may generate the one or more intermediate analytics results by initially identifying one or more candidate features required for determining the digital identity trust index of the subscriber, and one or more sub-features related to each of the one or more candidate features, from at least one of the secure sensitive contextual data 209 and the non-sensitive contextual data. In some embodiments, the one or more candidate features and the one or more sub-features related to each of the one or more candidate features may be determined based on an optimal feature set. In some embodiments, optimal feature set may be a subset of features of the secure sensitive contextual data 209, wherein the optimal feature set may include the most relevant and important features required for determining the digital identity trust index. In some embodiments, the optimal feature set is determined during a training phase of a machine learning model used in the trust index determining server 107. The optimal feature set is thereafter configured in each of the plurality of third party analytics agents 105 for performing the federated analytics. An exemplary optimal feature set for contextual data retrieved from each of the plurality of repositories 101 is as shown in the below Table 1. Sl.no Data Repository Optimal feature set
1. Financial transactions Repository 1. Credit payment history (late payments, payment defaults)
2. Current debt Burden
3. Business market value assessment
4. Business presence in multiple sectors
5. Financial transactions from multiple banks
2. Tax/Goods and Services Tax (GST) records Repository 1. GST payment defaults
2. Property tax
3. Tax bracket
3. Insurance data repository 1. Amount insured
2. Past insurance claims
3. Types of insurance (Financial/Health/Property)
4. Insurance payment defaults
4. Investment records repository 1. Investment portfolio
2. Premium payment defaults
5. Social Media data repository 1. Behavioural analysis
6. Demographic statistics repository 1. Business demand supply information
2. Employment/Financial status
3. Infrastructure Quality
4. Population density
7. Identity verification repository 1. National Identity verification
2. Business registration
3. E-verified documents
8. Health stack information repository 1. Customer health record categorization
2. Health service provider records
9. Police, Central Bureau of Investigation (CBI) and judiciary records repository 1. Criminal background status
10. Online customer feedback repository 1. Customer ratings
11. Service quality records repository 1. Service provider quality of services
12. Blockchain records repository 1. Land/Property records
2. Education certificates
3. E-driving license and Records
4. E-passport and Travel records
13. Telecom data repository 1. Location based service records Table 1 Further, the intermediate result generating module 227 may determine a normalized value with respect to a predefined scale, for each of the one or more candidate features using the one or more sub-features related to each of the one or more candidate features. In some embodiments, the one or more candidate features and the normalized value corresponding to each of the one or more candidate features constitute the one or more intermediate results. The one or more intermediate results are stored as the intermediate result data 211. As an example, if “financial transactions” is considered as the candidate feature 1 for determining the digital trust index, the one or more sub-features of the candidate feature “financial transaction” may be ? Date-Time: [Input to Candidate feature 1 Value] ? ATM/POS/Bank Location: [Input to Candidate feature 1 Value] ? Transaction Amount: [Input to Candidate feature 1 Value] Similarly, if “credit payment history and default” is considered as the candidate feature 2 for determining the digital trust index, the one or more sub-features of the candidate feature “credit payment history and default” may be ? Date-Time: [Input to Candidate feature 2 Value] ? Credit Amount: [Input to Candidate feature 2 Value] ? Annual Late Payments: [Input to Candidate feature 2 Value] ? Annual Payment Defaults [Input to Candidate feature 2 Value] For each of the one or more exemplary candidate features, the intermediate result generating module 227 may determine a normalized value with respect to a predefined scale using the one or more exemplary sub-features related to each of the one or more exemplary candidate features. As an example, the predefined scale may be between 0-100. In some embodiments, normalization is performed in order to convert general numeric values of the one or more candidate features into numeric values corresponding to the predefined scale. In some embodiments, machine learning models work more efficiently with normalized values than general numeric values. Further, in some embodiments, the transmitting module 229 may transmit each of the one or more intermediate analytics results to the trust index determining server 107 associated with the computing device 103. In some embodiments, till now the present disclosure has been described in view of one computing device 103, one data repository associated with the computing device 103 and one third party analytics agent associated with the computing device 103. The same procedure is followed for each of the plurality of computing devices 103 and the corresponding plurality of data repositories 101 and the corresponding plurality of third party analytics agents 105. Therefore, each of the plurality of computing devices 103 that received the data request from the trust index determining server 107, transmit each of the one or more intermediate analytics results to the trust index determining server 107, for determining a digital identity trust index in digital ecosystems. FIG.2D shows a detailed block diagram of a trust index determining server for determining a digital identity trust index in digital ecosystems in accordance with some embodiments of the present disclosure. In some implementations, the trust index determining server 107 may include data 235 and modules 237. As an example, the data 235 of the trust index determining server 107 is stored in the memory 119 configured in the trust index determining server 107 as shown in the FIG.2D. In one embodiment, the data 235 of the trust index determining server 107 may include intermediate result data 211, predicted data 239 and other data 241. In the illustrated FIG.2D, modules 237 of the trust index determining server 107 are described herein in detail. In some embodiments, the data 235 may be stored in the memory 119 in form of various data structures. Additionally, the data 235 can be organized using data models, such as relational or hierarchical data models. In some embodiments, the data 235 of the trust index determining server 107 may be stored in the memory 119 of the trust index determining server 107 in the form of various data structures. Additionally, the data 235 of the trust index determining server 107 can be organized using data models, such as relational or hierarchical data models. In some embodiments, the other data 235 of the trust index determining server 107 may store data, including temporary data and temporary files, generated by the modules 237 of the trust index determining server 107 for performing the various functions of the trust index determining server 107. In some embodiments, the data 235 of the trust index determining server 107 stored in the memory 119 may be processed by the modules 237 of the trust index determining server 107. The modules 237 of the trust index determining server 107 may be stored within the memory 119 of the trust index determining server 107. In an example, the modules 237 communicatively coupled to the processor 115 configured in the trust index determining server 107, may also be present outside the memory 119 as shown in FIG.2D and implemented as hardware. As used herein, the term modules refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or group) and memory that execute one or more software or firmware programs, a combinational logic circuit, and/or other suitable components that provide the described functionality. In an embodiment, the modules 237 of the trust index determining server 107 may include, for example, a receiving module 243, a feature selection module 245, a trust index determining module 247, a trust index updating module 249 and other modules 251. The other modules 251 may be used to perform various miscellaneous functionalities of the trust index determining server 107. It will be appreciated that such aforementioned modules 237 of the trust index determining server 107 may be represented as a single module or a combination of different modules. In some embodiments, the receiving module 243 may receive one or more intermediate analytics results from a plurality of computing devices 103 associated with a corresponding plurality of data repositories 101. In some embodiments, the one or more intermediate results are related to a subscriber whose digital identity trust index is to be determined. The one or more intermediate results may include, but not limited to, one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. Further, the feature selection module 245 may select a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set (This phase may be referred as prediction phase). In some embodiments, the key feature of the one or more candidate features, and the rest of the one or more candidate features are determined during a training phase of the machine learning model, based on historical data related to an application for which the digital identity trust index of the subscriber is being determined. In some embodiments, the historical data used for training the machine learning model is data collected for a large number of users. In some embodiments, supervised machine learning models may be used in the present disclosure. As an example, consider digital identity trust index is being determined for financial credit assessment application. During the training phase, the machine learning model may be trained using historical data features such as “credit payment history” as an output tag and rest of the features as an input tag, to predict the probability of credit payment, which would be considered as the digital identity trust index of the user for financial credit assessment application. As an example, the rest of the features for the financial credit assessment application may be “financial transactions from multiple banks”, “current debt burden”, “business market value assessment” and the like. Therefore, during the training phase, the machine learning model is trained by tagging the main feature whose prediction is required, with an output tag and rest of the features with an input tag. Referring back to the prediction phase, the trust index determining module 247 may generate a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features, as an input, to the machine learning model. In some embodiments, based on the training that the machine learning model has undergone during the training phase, the machine learning model may generate the prediction result of the key feature tagged as output, based on the optimal feature set. In some embodiments, the generated prediction result is a probability associated with the key feature tagged as “Output”, which is considered as the digital identity trust index of the subscriber. The prediction result generated by the trust index determining module 247 may be stored as the predicted data 239. As an example, in a business assessment application, “business continuity” may be selected as the key feature among the one or more candidate features, and therefore is tagged as “Output”. However, rest of the candidate features such as “business presence”, “business market value”, “business debts” and the like may be tagged as “Input”. Upon providing the rest of the candidate features to the machine learning model during the prediction phase, the machine learning model may predict the probability of business continuity, which is the probability of the key feature tagged as “Output”. The prediction result i.e. probability of business continuity may be considered as digital identity trust index of the subscriber for the application of business assessment. Further, in some embodiments, the trust index updating module 249 may determine an updated digital identity trust index, periodically to reflect changes in the information related to the subscriber. The periodicity may be pre-set by a service provider of the digital identity trust index. In some embodiments, the updated digital identity trust index may be determined by determining an exponential moving average of currently determined digital identity trust index based on updated contextual data, and the previously determined digital identity trust index. The current digital identity trust score may be determined using the below Equation 3. Current_Digital_Identity_Trust_Index = Digital_Identity_Trust_Index [n] * Exp_Percent + Digital_Identity_Trust_Index [n-1] * (1 - Exp_Percent) ------ Equation 3 In the above Equation 3, Current_Digital_Identity_Trust_Index denotes the current value of digital identity trust index after applying Exponential Moving Average. Digital_Identity_Trust_Index [n] denotes nth computation of the digital identity trust index, where n= 1, 2, 3 …. Digital_Identity_Trust_Index [n-1] denotes n-1th computation of the digital identity trust index, where n-1 = 1, 2, 3…. Time_Period denotes time duration between consecutive computations of the digital identity trust index, or in other words, periodicity set by the service provider for updating the digital identity trust index. Exp_Percent denotes a decay constant. In some embodiments, the trust index updating module 249 may determine the value of Exp_Percent based on the below Equation 4. Exp_Percent = 2 / (Time_Period + 1) ------------- Equation 4 Henceforth, the process of determining a digital identity trust index in a digital ecosystem is explained with the help of one or more examples for better understanding of the present disclosure. However, the one or more examples should not be considered as limitation of the present disclosure. Consider an exemplary application such as financial credit assessment application. In some embodiments, financial credit assessment application is just exemplary and should not be construed as a limitation of the present disclosure. The method disclosed in the present disclosure can be applicable for wide variety of applications such as employee on-boarding trust score, mobility service provider trust score, demographic trust score for smart city planning, data trust score for information sharing, logistics service provider trust score and the like. Consider the subscriber for whom the digital identity trust index has to be determined is a Small/Medium Entity (SME). Consider the subscriber 253 approaches a financial entity 255 for a loan. The financial entity 255 may avail the services of the trust index determining server 107 to determine the digital identity trust index of the subscriber 253 in real-time, as shown in the FIG.2E. Thereafter, the trust index determining server 107 may send a data request to exemplary data repositories such as financial transactions repository 1011, identity verification repository 1012 and police records repository 1013. Consider the computing devices 1031, 1032, and 1033 are associated with the financial transactions repository 1011, identity verification repository 1012 and police records repository 1013, respectively, as shown in the FIG.2F. Upon receiving the data request, each of the computing devices 1031, 1032, and 1033 may retrieve contextual data from the corresponding data repositories 1011, 1012, and 1013. As an example, consider the contextual data retrieved from the financial transactions repository 1011 comprises the following data of the subscriber 253 related to the financial transactions, as shown in the below Table 2. Also, the below Table 2 shows whether the retrieved contextual data is sensitive/non-sensitive, along with its degree of sensitivity. Sl.no Contextual data (also referred as candidate features) Contextual elements (also referred as sub-features of the candidate features) Sensitive Non-sensitive Degree of sensitivity
1. Financial transactions from multiple banks 1. Date and time Yes - Critical
2. ATM/POS/Bank location Yes - Critical
3. Transaction Amount Yes - High
2. Credit payment history and defaults 1. Date and time Yes - Medium
2. Credit Amount - Yes -
3. Annual late payments - Yes -
4. Annual payment defaults - Yes -
3. Current Debt Burden 1. No. of entities from which debt is already taken - Yes -
2. Debt burden amount - Yes -
4. Business market value assessment Assessment amount - Yes -
5. Business presence in multiple sectors Number of sectors - Yes - Table 2 Similarly, the contextual data may be retrieved from identity verification repository 1012 and police records repository 1013. Upon retrieving the contextual data, the computing devices 1031, 1032, and 1033 may generate secure sensitive contextual data 209 by applying noise perturbed feature based technique on the sensitive contextual data retrieved from the corresponding data repositories 1011, 1012, and 1013. Thereafter, the computing devices 1031, 1032, and 1033 may provide the corresponding non-sensitive contextual data and the secure sensitive contextual data 209 to the corresponding third party analytics agents 1051, 1052, and 1053, respectively, as shown in the FIG.2F. Each of the third party analytics agents 1051, 1052, and 1053 may generate one or more intermediate analytics results based on the one or more candidate features and the one or more sub-features related to each of the one or more candidate features, that are determined based on an optimal feature set. The one or more intermediate analytics results of the one or more candidate features and the one or more sub-features related to each of the one or more candidate features shown in the above Table 2, are given below: o Financial Transactions from Multiple Banks: [Candidate Feature 1 (CF 1) Value: On Scale of 100] ? Date-Time: [Input to CF 1 Value] ? ATM/POS/Bank Location: [Input to CF 1 Value] ? Transaction Amount: [Input to CF 1 Value] o Credit Payment History and Defaults: [CF 2 Value: On Scale of 100] ? Date-Time: [Input to CF 2 Value] ? Credit Amount: [Input to CF 2 Value] ? Annual Late Payments: [Input to CF 2 Value] ? Annual Payment Defaults [Input to CF 2 Value] o Current Debt Burden: [CF 3 Value: On Scale of 100] ? Number of Entities from which Credit is already taken: [Input to CF 3 Value] ? Debt Burden Amount: [Input to CF 3 Value] o Business Market Value Assessment: [CF 4 Value: On Scale of 100] ? Assessment Amount: [Input to CF 4 Value] o Business Presence in Multiple Sectors: [CF 5 Value: On Scale of 100] ? Number of Sectors (Finance/Travel/Ecommerce): [Input to CF 5 Value] Similarly, the one or more intermediate analytics results may be determined for the one or more candidate features and the one or more sub-features related to the one or more candidate features of the contextual data retrieved from identity verification repository 1012 and police records repository 1013. The computing devices 1031, 1032, and 1033 may transmit the corresponding one or more intermediate analytics results to the trust index determining server 107. Among the one or more intermediate analytics results, the trust index determining server 107 may select a key feature among the one or more candidate features, based on the optimal feature set. Since, the application is a financial credit assessment application, the trust index determining server 107 selects “Credit payment history and defaults” as the key feature among the one or more candidate features, based on the optimal feature set. Therefore, the key feature and rest of the one or more candidate features are tagged as shown in the below Table 3. Candidate feature type Candidate feature Tag assigned to the candidate feature Key feature Credit payment history and defaults Output Rest of the one or more candidate features Financial transactions from multiple banks Input Current Debt Burden Input Business market value assessment Input Business presence in multiple sectors . . . Input Table 3 Further, the trust index determining server 107 may predict a probability associated with the key feature tagged as “Output”, i.e. probability of the credit payment, by providing the rest of the one or more candidate features tagged as “Input” to the machine learning model of the trust index determining server 107. The prediction result i.e. the numeric value thus generated for probability of the credit payment, is considered as the digital identity trust index of the subscriber for the financial credit assessment application. FIG.3A shows flowchart illustrating a method of generating analytics results for determining a digital identity trust index in digital ecosystem, in accordance with some embodiments of the present disclosure. As illustrated in FIG.3A, the method 300a includes one or more blocks illustrating a method of generating analytics results for determining a digital identity trust index in digital ecosystem. The method 300a may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform functions or implement abstract data types. The order in which the method 300a is described is not intended to be construed as a limitation, and any number of the described method blocks can be combined in any order to implement the method 300a. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described herein. Furthermore, the method 300a can be implemented in any suitable hardware, software, firmware, or combination thereof. At block 301, the method includes receiving, by a processor 109 of the computing device 103, a data request from a trust index determining server 107. In some embodiments, the data request is related to a subscriber such as a customer, a service provider or a small/medium enterprise. At block 303, the method includes retrieving, by the processor 109 of the computing device 103, contextual data related to the subscriber, from a corresponding data repository of the plurality of data repositories 101. In some embodiments, the contextual data may include at least one of sensitive contextual data and non-sensitive contextual data. At block 305, the method includes generating, by the processor 109 of the computing device 103, secure sensitive contextual data 209 by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique. In some embodiments, the degree of sensitivity of each contextual element may be quantified into one of “Critical”, “High”, “Medium” and “Low”. At block 307, the method includes generating, by the processor 109 of the computing device 103, one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data 209 and the non-sensitive contextual data, using a third party analytics agent 105 associated with the computing device 103. In some embodiments, the processor 109 of the computing device 103 may transmit the one or more intermediate analytics results to the trust index determining server 107, for determining a digital identity trust index in digital ecosystems. FIG.3B shows flowchart illustrating a method of determining a digital identity trust index in digital ecosystems, in accordance with some embodiments of the present disclosure. As illustrated in FIG.3B, the method 300b includes one or more blocks illustrating a method of determining a digital identity trust index in digital ecosystems. The method 300b may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform functions or implement abstract data types. The order in which the method 300b is described is not intended to be construed as a limitation, and any number of the described method blocks can be combined in any order to implement the method 300b. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described herein. Furthermore, the method 300b can be implemented in any suitable hardware, software, firmware, or combination thereof. At block 311, the method includes receiving, by a processor 115 of a trust index determining server 107, one or more intermediate analytics results from a plurality of computing devices 103 associated with a corresponding plurality of data repositories 101. In some embodiments, the one or more intermediate results are related to a subscriber whose digital identity trust index is to be determined. In some embodiments, the one or more intermediate results may include, one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. At block 313, the method includes selecting, by the processor 115 of the trust index determining server 107, a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set. In some embodiments, the optimal feature set, the key feature of the one or more candidate features, and the rest of the one or more candidate features may be determined during a training phase of the machine learning model, based on historical data related to an application for which the digital identity trust index of the subscriber is being determined. At block 315, the method includes generating, by the processor 115 of the trust index determining server 107, a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber. In some embodiments, the machine learning model may be trained to generate the prediction result based on the optimal feature set. In some embodiments, the processor 115 of the trust index determining server 107 may determine an updated digital identity trust index, periodically, by determining an exponential moving average of currently determined digital identity trust index based on updated contextual data, and the previously determined digital identity trust index. FIG.4 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure. In some embodiments, FIG.4 illustrates a block diagram of an exemplary computer system 400 for implementing embodiments consistent with the present invention. In some embodiments, the computer system 400 can be a computing device 103 used for generating analytics results for determining a digital identity trust index in digital ecosystem. In some embodiments, the computer system 400 can be a trust index determining server 107 used for determining a digital identity trust index in digital ecosystems. FIG.4 is illustrated considering the computer system 400 as the computing device 103. However, this should not be considered as a limitation, since the computer system 400 can be the trust index determining server 107 as well (not shown in FIG.4). The computer system 400 may include a central processing unit (“CPU” or “processor”) 402. The processor 402 may include at least one data processor for executing program components for executing user or system-generated business processes. A user may include a person, a person using a device such as those included in this invention, or such a device itself. The processor 402 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc. The processor 402 may be disposed in communication with input devices 411 and output devices 412 via I/O interface 401. The I/O interface 401 may employ communication protocols/methods such as, without limitation, audio, analog, digital, stereo, IEEE-1394, serial bus, Universal Serial Bus (USB), infrared, PS/2, BNC, coaxial, component, composite, Digital Visual Interface (DVI), high-definition multimedia interface (HDMI), Radio Frequency (RF) antennas, S-Video, Video Graphics Array (VGA), IEEE 802.n /b/g/n/x, Bluetooth, cellular (e.g., Code-Division Multiple Access (CDMA), High-Speed Packet Access (HSPA+), Global System For Mobile Communications (GSM), Long-Term Evolution (LTE), WiMax, or the like), etc. Using the I/O interface 401, computer system 400 may communicate with input devices 411 and output devices 412. In some embodiments, the processor 402 may be disposed in communication with a communication network 409 via a network interface 403. The network interface 403 may communicate with the communication network 409. The network interface 403 may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10/100/1000 Base T), Transmission Control Protocol/Internet Protocol (TCP/IP), token ring, IEEE 802.11a/b/g/n/x, etc. Using the network interface 403 and the communication network 409, the computer system 400 may communicate with plurality of data repositories 101 (1011 up to 101n), plurality of third party analytics agents 105 (1051 up to 105n) and a trust index determining server 107. In some embodiments, the computer system 400 may communicate via IP based communication and non-internet or non-IP based communication such as Universal Serial Bus (USB), Bluetooth and the like. The communication network 409 can be implemented as one of the different types of networks, such as intranet or Local Area Network (LAN), Closed Area Network (CAN) and such within the autonomous vehicle. The communication network 409 may either be a dedicated network or a shared network, which represents an association of the different types of networks that use a variety of protocols, for example, Hypertext Transfer Protocol (HTTP), CAN Protocol, Transmission Control Protocol/Internet Protocol (TCP/IP), Wireless Application Protocol (WAP), etc., to communicate with each other. Further, the communication network 409 may include a variety of network devices, including routers, bridges, servers, computing devices, storage devices, etc. In some embodiments, the processor 402 may be disposed in communication with a memory 405 (e.g., RAM, ROM, etc. not shown in FIG.4) via a storage interface 404. The storage interface 404 may connect to memory 405 including, without limitation, memory drives, removable disc drives, etc., employing connection protocols such as Serial Advanced Technology Attachment (SATA), Integrated Drive Electronics (IDE), IEEE-1394, Universal Serial Bus (USB), fibre channel, Small Computer Systems Interface (SCSI), etc. The memory drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory devices, solid-state drives, etc. The memory 405 may store a collection of program or database components, including, without limitation, a user interface 406, an operating system 407, a web browser 408 etc. In some embodiments, the computer system 400 may store user/application data, such as the data, variables, records, etc. as described in this invention. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle or Sybase. The operating system 407 may facilitate resource management and operation of the computer system 400. Examples of operating systems 407 include, without limitation, APPLE® MACINTOSH® OS X®, UNIX®, UNIX-like system distributions (E.G., BERKELEY SOFTWARE DISTRIBUTION® (BSD), FREEBSD®, NETBSD®, OPENBSD, etc.), LINUX® DISTRIBUTIONS (E.G., RED HAT®, UBUNTU®, KUBUNTU®, etc.), IBM®OS/2®, MICROSOFT® WINDOWS® (XP®, VISTA®/7/8, 10 etc.), APPLE® IOS®, GOOGLETM ANDROIDTM, BLACKBERRY® OS, or the like. The User interface 406 may facilitate display, execution, interaction, manipulation, or operation of program components through textual or graphical facilities. For example, user interfaces 406 may provide computer interaction interface elements on a display system operatively connected to the computer system 400, such as cursors, icons, checkboxes, menus, scrollers, windows, widgets, etc. Graphical User Interfaces (GUIs) may be employed, including, without limitation, Apple® Macintosh® operating systems’ Aqua®, IBM® OS/2®, Microsoft® Windows® (e.g., Aero, Metro, etc.), web interface libraries (e.g., ActiveX®, Java®, Javascript®, AJAX, HTML, Adobe® Flash®, etc.), or the like. In some embodiments, the computer system 400 may implement the web browser 408 stored program components. The web browser 408 may be a hypertext viewing application, such as MICROSOFT® INTERNET EXPLORER®, GOOGLETM CHROMETM, MOZILLA® FIREFOX®, APPLE® SAFARI®, etc. Secure web browsing may be provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. Web browsers 408 may utilize facilities such as AJAX, DHTML, ADOBE® FLASH®, JAVASCRIPT®, JAVA®, Application Programming Interfaces (APIs), etc. In some embodiments, the computer system 400 may implement a mail server stored program component. The mail server may be an Internet mail server such as Microsoft Exchange, or the like. The mail server may utilize facilities such as Active Server Pages (ASP), ACTIVEX®, ANSI® C++/C#, MICROSOFT®, .NET, CGI SCRIPTS, JAVA®, JAVASCRIPT®, PERL®, PHP, PYTHON®, WEBOBJECTS®, etc. The mail server may utilize communication protocols such as Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), MICROSOFT® exchange, Post Office Protocol (POP), Simple Mail Transfer Protocol (SMTP), or the like. In some embodiments, the computer system 400 may implement a mail client stored program component. The mail client may be a mail viewing application, such as APPLE® MAIL, MICROSOFT® ENTOURAGE®, MICROSOFT® OUTLOOK®, MOZILLA® THUNDERBIRD®, etc. Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present invention. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., non-transitory. Examples include Random Access Memory (RAM), Read-Only Memory (ROM), volatile memory, non-volatile memory, hard drives, Compact Disc (CD) ROMs, Digital Video Disc (DVDs), flash drives, disks, and any other known physical storage media. Advantages of the present disclosure are illustrated herein. The present disclosure associates a digital identity trust index with subscribers such as customers, service providers and SMEs based on contextual interpretation of data records retrieved from wide variety of data repositories, thereby providing the advantages of: • Easy interpretation of digital trust level of the subscriber. • Increased scope of verification. • Accurate comprehension of digital trust index. • Reduction in fraud cases. • Early risk management. The present disclosure secures the sensitive contextual data using a noise perturbed feature based technique, which is applied by the computing devices, and generates intermediate results using third party analytics agents that are locally configured in the location of the corresponding data repositories, thereby providing the advantages of: • Maintaining privacy and confidentiality of the user data. • Maintaining user data in a locked-down location by minimizing data movement across the communication network, which ensures data security, and also conserves bandwidth as only intermediate results are transmitted to the trust index determining server. The present disclosure generates the prediction result by using a machine learning model which is trained to generate the prediction result based on an optimal feature set, thereby providing the advantages of: • Enabling real-time digital trust services, by reducing the amount of data to be processed for the prediction, which in turn reduces the processing complexity and enhances computation speed. • Achieving complete automation of digital identity trust index determination, which results in accurate digital trust comprehension as compared to conventional human interpretation of records. Overall, the present disclosure determines the digital identity trust index for wide variety of applications, accurately and in real-time, while maintaining user privacy and data security. A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary a variety of optional components are described to illustrate the wide variety of possible embodiments of the invention. When a single device or article is described herein, it will be apparent that more than one device/article (whether or not they cooperate) may be used in place of a single device/article. Similarly, where more than one device or article is described herein (whether or not they cooperate), it will be apparent that a single device/article may be used in place of the more than one device or article or a different number of devices/articles may be used instead of the shown number of devices or programs. The functionality and/or the features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality/features. Thus, other embodiments of the invention need not include the device itself. The specification has described a method and a system for determining digital identity trust index in digital ecosystems. The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that on-going technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein. Such alternatives fall within the scope of the disclosed embodiments. Also, the words "comprising," "having," "containing," and "including," and other similar forms are intended to be equivalent in meaning and be open-ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items, or meant to be limited to only the listed item or items. It must also be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based here on. Accordingly, the embodiments of the present invention are intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims. Referral numerals Reference Number Description 100 Architecture 101 Plurality of data repositories 103 Plurality of computing devices 105 Plurality of third party analytics agents 107 Trust index determining server 109 Processor of the computing device 111 I/O interface of the computing device 113 Memory of the computing device 203 Data of the computing device 205 Modules of the computing device 207 Contextual data 209 Secure sensitive contextual data 211 Intermediate result data 213 Other data 221 Receiving module of the computing device 223 Data retrieving module 225 Secure data generating module 227 Intermediate result generating module 229 Transmitting module 231 Other modules of the computing device 235 Data of the trust index determining server 237 Modules of the trust index determining server 239 Predicted data 241 Other data of the trust index determining server 243 Receiving module of the trust index determining server 245 Feature selection module 247 Trust index determining module 249 Trust index updating module 251 Other modules of the trust index determining server 253 Exemplary subscriber 255 Exemplary financial entity 400 Exemplary computer system 401 I/O Interface of the exemplary computer system 402 Processor of the exemplary computer system 403 Network interface 404 Storage interface 405 Memory of the exemplary computer system 406 User interface 407 Operating system 408 Web browser 409 Communication network 411 Input devices 412 Output devices
Claims:We claim:
1. A method of generating analytics results for determining a digital identity trust index in digital ecosystem, wherein the digital ecosystem comprises a plurality of computing devices (103) associated with a corresponding plurality of data repositories (101), the method comprising:
receiving, by a computing device (103), a data request from a trust index determining server (107), wherein the data request is related to a subscriber;
retrieving, by the computing device (103), contextual data (207) related to the subscriber, from a corresponding data repository (101) of the plurality of data repositories (101), wherein the contextual data (207) comprises at least one of sensitive contextual data and non-sensitive contextual data;
generating, by the computing device (103), secure sensitive contextual data (209) by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique; and
generating, by the computing device (103), one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data (209) and the non-sensitive contextual data, using a third party analytics agent (105) associated with the computing device (103), wherein the one or more intermediate analytics results are transmitted to the trust index determining server (107), for determining a digital identity trust index in digital ecosystems.
2. The method as claimed in claim 1, wherein the degree of sensitivity of each contextual element is determined by:
determining a predefined category of each of the contextual element, wherein each predefined category is associated with a degree of sensitivity; and
assigning each contextual element with the degree of sensitivity of the corresponding predefined category.
3. The method as claimed in claim 1, wherein the degree of sensitivity of each contextual element is quantified into one of “Critical”, “High”, “Medium” and “Low”.
4. The method as claimed in claim 1, wherein the noise perturbed feature based technique comprises:
dividing at least one of the sensitive contextual data and the non-sensitive contextual data into plurality of matrices, based on total number of the contextual elements of at least one of the sensitive contextual data and the non-sensitive contextual data, wherein the contextual elements of the sensitive contextual data are referred as sensitive contextual elements;
clustering a plurality of sensitive contextual elements among total number of sensitive contextual elements, into a plurality of clusters, based on a predefined clustering technique, wherein a centroid value is computed for each of the plurality of clusters; and
replacing a value of,
each of the plurality of sensitive contextual elements with a combined value obtained by combining the centroid value of each of the corresponding plurality of clusters and a first corresponding random noise, wherein the first random noise is selected for each of the plurality of sensitive contextual elements from a Laplace distribution, based on a degree of sensitivity of each of the plurality of sensitive contextual elements; and
each of rest of the plurality of sensitive contextual elements among the total number of sensitive contextual elements with a combined value obtained by combining original value of each corresponding sensitive contextual element and a second corresponding random noise, wherein the second random noise is selected for each of rest of the plurality of sensitive contextual elements from the Laplace distribution, based on a degree of sensitivity of each of the rest of the plurality of sensitive contextual elements.
5. The method as claimed in claim 1, wherein the subscriber is one of a customer, a service provider or a small/medium enterprise.
6. The method as claimed in claim 1, wherein generating the one or more intermediate analytics results by performing federated analytics comprises:
identifying, one or more candidate features required for determining the digital identity trust index of the subscriber, and one or more sub-features related to each of the one or more candidate features, from at least one of the secure sensitive contextual data (209) and the non-sensitive contextual data, based on an optimal feature set; and
determining a normalized value with respect to a predefined scale, for each of the one or more candidate features using the one or more sub-features related to each of the one or more candidate features, wherein the one or more candidate features and the normalized value corresponding to each of the one or more candidate features constitute the one or more intermediate results.
7. A method of determining a digital identity trust index in digital ecosystems, the method comprising:
receiving, by a trust index determining server (107), one or more intermediate analytics results from a plurality of computing devices (103) associated with a corresponding plurality of data repositories (101), wherein the one or more intermediate results are related to a subscriber whose digital identity trust index is to be determined, wherein the one or more intermediate results comprises one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features;
selecting, by the trust index determining server (107), a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set; and
generating, by the trust index determining server (107), a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber, wherein the machine learning model is trained to generate the prediction result based on the optimal feature set.
8. The method as claimed in claim 7, wherein the optimal feature set, the key feature of the one or more candidate features, and the rest of the one or more candidate features are determined during a training phase of the machine learning model, based on historical data related to an application for which the digital identity trust index of the subscriber is being determined.
9. The method as claimed in claim 7, wherein the received one or more intermediate analytics results are generated by performing federated analytics on at least one of secure sensitive contextual data (209) and non-sensitive contextual data, related to the subscriber, wherein the secure sensitive contextual data (209) is obtained by applying a noise perturbed feature based technique on sensitive contextual data related to the subscriber.
10. The method as claimed in claim 7 further comprises determining, by the trust index determining server (107), an updated digital identity trust index, periodically, by determining an exponential moving average of currently determined digital identity trust index based on updated contextual data, and the previously determined digital identity trust index.
11. The method as claimed in claim 7, wherein the subscriber is one of a customer, a service provider or a small/medium enterprise.
12. A computing device (103) for generating analytics results for determining a digital identity trust index in digital ecosystem, wherein the digital ecosystem comprises a plurality of computing devices (103) associated with a corresponding plurality of data repositories (101), the computing device (103) comprising:
a processor (109); and
a memory (113) communicatively coupled to the processor (109), wherein the memory (113) stores the processor (109)-executable instructions, which, on execution, causes the processor (109) to:
receive a data request from a trust index determining server, wherein the data request is related to a subscriber;
retrieve contextual data (207) related to the subscriber, from a corresponding data repository of the plurality of data repositories (101), wherein the contextual data (207) comprises at least one of sensitive contextual data and non-sensitive contextual data;
generate secure sensitive contextual data (209) by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique; and
generate one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data (209) and the non-sensitive contextual data, using a third party analytics agent (105) associated with the computing device (103), wherein the one or more intermediate analytics results are transmitted to the trust index determining server, for determining a digital identity trust index in digital ecosystems.
13. The computing device (103) as claimed in claim 12, wherein the processor (109) determines degree of sensitivity of each contextual element by:
determining a predefined category of each of the contextual element, wherein each predefined category is associated with a degree of sensitivity; and
assigning each contextual element with the degree of sensitivity of the corresponding predefined category.
14. The computing device (103) as claimed in claim 12, wherein the degree of sensitivity of each contextual element is quantified into one of “Critical”, “High”, “Medium” and “Low”.
15. The computing device (103) as claimed in claim 12, wherein to perform the noise perturbed feature based technique, the processor (109) is configured to:
divide at least one of the sensitive contextual data and the non-sensitive contextual data into plurality of matrices, based on total number of the contextual elements of at least one of the sensitive contextual data and the non-sensitive contextual data, wherein the contextual elements of the sensitive contextual data are referred as sensitive contextual elements;
cluster a plurality of sensitive contextual elements among total number of sensitive contextual elements, into a plurality of clusters, based on a predefined clustering technique, wherein a centroid value is computed for each of the plurality of clusters; and
replace a value of,
each of the plurality of sensitive contextual elements with a combined value obtained by combining the centroid value of each of the corresponding plurality of clusters and a first corresponding random noise, wherein the first random noise is selected for each of the plurality of sensitive contextual elements from a Laplace distribution, based on a degree of sensitivity of each of the plurality of sensitive contextual elements; and
each of rest of the plurality of sensitive contextual elements among the total number of sensitive contextual elements with a combined value obtained by combining original value of each corresponding sensitive contextual element and a second corresponding random noise, wherein the second random noise is selected for each of rest of the plurality of sensitive contextual elements from the Laplace distribution, based on a degree of sensitivity of each of the rest of the plurality of sensitive contextual elements.
16. The computing device (103) as claimed in claim 12, wherein the subscriber is one of a customer, a service provider or a small/medium enterprise.
17. The computing device (103) as claimed in claim 12, wherein to generate the one or more intermediate analytics results by performing federated analytics, the processor (109) is configured to:
identify one or more candidate features required for determining the digital identity trust index of the subscriber, and one or more sub-features related to each of the one or more candidate features, from at least one of the secure sensitive contextual data (209) and the non-sensitive contextual data, based on an optimal feature set; and
determine a normalized value with respect to a predefined scale, for each of the one or more candidate features using the one or more sub-features related to each of the one or more candidate features, wherein the one or more candidate features and the normalized value corresponding to each of the one or more candidate features constitute the one or more intermediate results.
18. A trust index determining server (107) for determining a digital identity trust index in digital ecosystems, the trust index determining server (107) comprising:
a processor (115); and
a memory (119) communicatively coupled to the processor (115), wherein the memory (119) stores the processor-executable instructions, which, on execution, causes the processor (115) to:
receive one or more intermediate analytics results from a plurality of computing devices (103) associated with a corresponding plurality of data repositories (101), wherein the one or more intermediate results are related to a subscriber, whose digital identity trust index is to be determined, wherein the one or more intermediate results comprises one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features;
select a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set; and
generate a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber, wherein the machine learning model is trained to generate the prediction result based on the optimal feature set.
19. The trust index determining server (107) as claimed in claim 18, wherein the optimal feature set, the key feature of the one or more candidate features, and the rest of the one or more candidate features are determined during a training phase of the machine learning model, based on historical data related to an application for which the digital identity trust index of the subscriber is being determined.
20. The trust index determining server (107) as claimed in claim 18, wherein the received one or more intermediate analytics results are generated by performing federated analytics on at least one of secure sensitive contextual data (209) and non-sensitive contextual data, related to the subscriber, wherein the secure sensitive contextual data (209) is obtained by applying a noise perturbed feature based technique on sensitive contextual data related to the subscriber.
21. The trust index determining server (107) as claimed in claim 18, wherein the processor (115) is further configured to determine an updated digital identity trust index, periodically, by determining an exponential moving average of currently determined digital identity trust index based on updated contextual data, and the previously determined digital identity trust index.
22. The trust index determining server (107) as claimed in claim 18, wherein the subscriber is one of a customer, a service provider or a small/medium enterprise.
, Description:TECHNICAL FIELD
The present subject matter is related, in general to digital trust and more particularly, but not exclusively to method and system for determining digital identity trust index in digital ecosystems.
BACKGROUND
In today’s digital world, rate of economic growth depends on emergence of new business opportunities and level of trust among participating entities. Generation of huge amount of digital data has led to creation of diverse public and private data repositories and platforms that support building business applications providing rich insights by data analytics. However, for co-creation and sustained usage of new data driven applications, and for digital transformation of economies, there exists a need to build trust among the participating entities, such as customers, service providers and Small/Medium Enterprises (SMEs). Existing verification processes are built on submission of physical documents, which have a limited scope with respect to data sources. Moreover, verification based on submitted physical documents is time-consuming, unreliable and based on human interpretation of data records, which leads to high number of fraud cases, poor services and reduced digital trust.
Currently, for ensuring digital trust, scores such as Credit Information Bureau (India) Limited (CIBIL®) and Zhima® credit associate a number to a customer, which defines level of trust and credit worthiness of the customer. Banks and other organizations provide incentives to the customer based on the score. However, the CIBIL® and Zhima® credit scores are limited only to financial industry, due to which, CIBIL® score relies only on financial transaction history of the customer as source data, and Zhima® credit score relies on financial transaction history as well as behavioural analysis of the customers, while using specific online websites, as the source data.
Apart from CIBIL® and Zhima® credit scores mentioned above, currently there are other techniques in use for computing identity and trust worthiness of online users. Such techniques integrate user data retrieved from social media, e-commerce and other public repositories along with subscriber data provided directly by the user. Upon retrieving the user data, these techniques compare correctness of subscriber data values with the user data retrieved from limited number of public repositories, to determine trust score for the user. However, such techniques are completely dependent on correctness of the subscriber data provided by the user rather than contextual interpretation of user data retrieved from public repositories, which leads to inaccuracy in trust score determination. Moreover, such techniques retrieve user data from limited number of public repositories, which may lead to the risk of missing out on important information related to the user.
Yet another existing technique allows a user to select a service provider from a group, based on trust scores associated with each of the service providers. In this existing technique, the service provider trust score is determined based on user reputation score and a general reputation score. User reputation score is a score computed based on current user’s perception of the service provider, and the general reputation score is a score computed based on generic perception of other users who have already received services of the service provider. Therefore, this technique completely relies on the perception of users to determine the service provider trust score, which is just one dimension for determining the trust score. Such single dimensional computations have an extremely narrow scope, and are inaccurate due to their sole dependency on perception of users, rather than user data and facts available in public repositories.
Therefore, there exists a need to accurately determine the digital identity trust index, for wide variety of applications in digital ecosystems.
The information disclosed in this background of the disclosure section is only for enhancement of understanding of the general background of the invention and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art.
SUMMARY
Disclosed herein is a method of generating analytics results for determining a digital identity trust index in digital ecosystem. The digital ecosystem comprises a plurality of computing devices associated with a corresponding plurality of data repositories. The method includes receiving, by a computing device, a data request from a trust index determining server, wherein the data request is related to a subscriber. Further, the method includes retrieving contextual data related to the subscriber, from a corresponding data repository of the plurality of data repositories. The contextual data comprises at least one of sensitive contextual data and non-sensitive contextual data. Subsequently, the method includes generating secure sensitive contextual data by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique. Upon generating the secure sensitive contextual data, the method includes generating one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data and the non-sensitive contextual data, using a third party analytics agent associated with the computing device. The one or more intermediate analytics results are transmitted to the trust index determining server, for determining a digital identity trust index in digital ecosystems.
Further, the present disclosure discloses a method of determining a digital identity trust index in digital ecosystems. The method includes receiving, by a trust index determining server, one or more intermediate analytics results from a plurality of computing devices associated with a corresponding plurality of data repositories. The one or more intermediate results are related to a subscriber whose digital identity trust index is to be determined. The one or more intermediate results include one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. Further, the method includes selecting a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set. Finally, the method includes generating a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber. The machine learning model is trained to generate the prediction result based on the optimal feature set.
Further, the present disclosure includes a computing device for generating analytics results for determining a digital identity trust index in digital ecosystem. The digital ecosystem includes a plurality of computing devices associated with a corresponding plurality of data repositories. The computing device includes a processor and a memory communicatively coupled to the processor. The memory stores the processor-executable instructions, which, on execution, causes the processor to receive a data request from a trust index determining server, wherein the data request is related to a subscriber. Further, the processor retrieves contextual data related to the subscriber, from a corresponding data repository of the plurality of data repositories. The contextual data comprises at least one of sensitive contextual data and non-sensitive contextual data. Further, the processor generates secure sensitive contextual data by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique. Upon generating the secure sensitive contextual data, the processor generates one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data and the non-sensitive contextual data, using a third party analytics agent associated with the computing device. The one or more intermediate analytics results are transmitted to the trust index determining server, for determining a digital identity trust index in digital ecosystems.
Furthermore, the present disclosure discloses a trust index determining server for determining a digital identity trust index in digital ecosystems. The trust index determining server comprises a processor and a memory communicatively coupled to the processor. The memory stores the processor-executable instructions, which, on execution, causes the processor to receive one or more intermediate analytics results from a plurality of computing devices associated with a corresponding plurality of data repositories. The one or more intermediate results are related to a subscriber, whose digital identity trust index is to be determined. The one or more intermediate results comprises one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. Thereafter, the processor selects a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set. Finally, the processor generates a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber. The machine learning model is trained to generate the prediction result based on the optimal feature set.
The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.
BRIEF DESCRIPTION OF THE ACCOMPANYING DIAGRAMS
The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The same numbers are used throughout the figures to reference like features and components. Some embodiments of system and/or methods in accordance with embodiments of the present subject matter are now described, by way of example only, and with reference to the accompanying figures, in which:
FIG.1A shows an exemplary architecture for generating analytics results and determining a digital identity trust index in a digital ecosystem in accordance with some embodiments of the present disclosure;
FIG.1B shows a brief block diagram of a computing device in accordance with some embodiments of the present disclosure;
FIG.1C shows a brief block diagram of a trust index determining server in accordance with some embodiments of the present disclosure;
FIG.2A shows a detailed block diagram of a computing device for generating analytics results for determining a digital identity trust index in digital ecosystem in accordance with some embodiments of the present disclosure;
FIG.2B shows an exemplary matrix comprising sensitive and non-sensitive data in accordance with some embodiments of the present disclosure;
FIG.2C shows an exemplary matrix with clusters in accordance with some embodiments of the present disclosure;
FIG.2D shows a detailed block diagram of a trust index determining server for determining a digital identity trust index in digital ecosystems in accordance with some embodiments of the present disclosure;
FIG.2E and FIG.2F show exemplary scenarios for determining a digital identity trust index in a digital ecosystem in accordance with some embodiments of the present disclosure;
FIG.3A shows a flowchart illustrating a method of generating analytics results for determining a digital identity trust index in digital ecosystem in accordance with some embodiments of the present disclosure;
FIG.3B shows a flowchart illustrating a method of determining a digital identity trust index in digital ecosystems in accordance with some embodiments of the present disclosure; and
FIG.4 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.
It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown.
DETAILED DESCRIPTION
In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.
While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the scope of the disclosure.
The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises… a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or method.
Disclosed herein are method and apparatus for determining digital identity trust index in digital ecosystems for a wide variety of applications. The method includes receiving, by a computing device, a data request from a trust index determining server. In some embodiments, the data request may be related to a subscriber. In some embodiments, the subscriber may be of a customer, a service provider or a small/medium enterprise. Further, the computing device may retrieve contextual data related to the subscriber, from a corresponding data repository of the plurality of data repositories. As an example, the contextual data may be financial data, tax related data, identity verification data, police and judicial data, social media data, health related data and the like, related to the subscriber. The contextual data may include at least one of sensitive contextual data and non-sensitive contextual data. In the present disclosure, the computing devices retrieve the contextual data from the corresponding data repositories and provide the contextual data to the corresponding third party analytics agents. In some embodiments, in order to secure privacy of the sensitive contextual data related to the subscriber, the present disclosure employs a noise perturbed feature based technique to add noise to the sensitive contextual data, prior to providing the sensitive contextual data to the corresponding third party analytics agents. The third party analytics agents may perform federated analytics on the non-sensitive contextual data and secure sensitive contextual data, which results in one or more intermediate analytics results.
The one or more intermediate results may include, but not limited to, one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. In some embodiments, the plurality of computing devices may transmit the one or more intermediate results to a trust index determining server for further processing. The trust index determining server may generate a prediction result based on each of the one or more intermediate results, for determining the digital identity trust index of the subscriber. In some embodiments, the prediction result may be generated using a machine learning model which is trained to generate the prediction result based on an optimal feature set.
The present disclosure allows determination of digital trust index for wide variety of applications such as merchant trust score, data trust score for information sharing, customer trust score for financial services, logistics service provider trust score, demographic trust score for smart city planning, e-commerce service provider trust score, employee onboarding trust score, mobility service provider trust score and the like. Further, since the present disclosure determines the digital identity trust index by performing contextual interpretation of the user data retrieved from wide variety of public and private data repositories, the scope of verification is broad, which in turn results in accurate comprehension of trust levels of the customer, service provider or the small/medium enterprises. Such accurate comprehension of trust levels reduces the number of fraud cases and provides early risk management to businesses. Further, the noise perturbed feature based technique used in the present disclosure anonymizes the sensitive contextual data to an extent that it ensures confidentiality of the sensitive contextual data, but also allows to perform meaningful data analytics, by the third party analytics agents. Moreover, since the present disclosure deploys the third party analytics agents at the respective location of each of the public and private data repositories, contextual data of the user remains in its locked-down location, and intermediate analytics results are generated locally. This eliminates the need to transmit large amount of the contextual data over the communication network, thereby preventing exposure of the contextual data to security breaches and data theft. Moreover, transmitting only the intermediate analytics results to the trust index determining server eliminates over utilization of the available bandwidth, thereby enabling bandwidth conservation. Further, since the present disclosure generates prediction result based on optimal feature set, the present disclosure reduces processing complexity and enhances the speed of determination of the digital trust identity index, which helps in performing/providing services dependent on digital identity trust index of the subscriber, in real-time.. Moreover, the machine learning model used in the present disclosure automates the process of determining digital identity trust index, which means human interpretation of records is eliminated, thereby improving accuracy of the digital identity trust index. Overall, the present disclosure determines the digital identity trust index for wide variety of applications, accurately and in real-time, while maintaining user privacy and data security.
In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense.
FIG.1A shows an exemplary architecture for generating analytics results and determining a digital identity trust index in a digital ecosystem in accordance with some embodiments of the present disclosure.
The architecture 100 comprises a data repository 1 1011 to a data repository n 101n (collectively referred to as plurality of data repositories 101), a computing device 1 1031 to a computing device n 103n (collectively referred to as plurality of computing devices 103), a third party analytics agent 1 1051 to a third party analytics agent n 105n (collectively referred to as plurality of third party analytics agents 105), and a trust index determining server 107. The plurality of data repositories 101 may be public data repositories or private repositories. As an example, the plurality of data repositories 101 may include, but not limited to, financial transactions repository, Tax/Goods and Services Tax (GST) records repository, insurance data repository, investment records repository, social media data repository, demographic statistics repository, identity verification repository, health stack information repository, police records repository, Central Bureau of Investigation (CBI) records repository, judiciary records repository, online customer feedback repository, service quality records repository, blockchain records repository and telecom data repository. In some embodiments, each of the plurality of data repositories 101 may be associated with each of the corresponding plurality of computing devices 103 as shown in the FIG.1A. In some other embodiments, the plurality of data repositories 101 may be configured within the corresponding plurality of computing devices 103. As an example, the plurality of computing devices 103 may include, but not limited to, a laptop, a computer, a desktop, a mobile phone, a tablet phone and the like. Further, each of the plurality of computing devices 103 may be associated with the corresponding plurality of third party analytics agents 105 as shown in the FIG.1A. In some embodiments, the plurality of third party analytics agents 105 may be configured within the corresponding plurality of computing devices 103. Each of the plurality of third party analytics agents 105 are locally configured i.e. configured at the same location as that of the corresponding plurality of computing devices 103. In some embodiments, each of the plurality of computing devices 103 may be associated with the trust index determining server 107 via a communication network (not shown in the FIG.1A). The communication network may be a wireless communication network. In some embodiments, the trust index determining server 107 may be configured remotely. In some other embodiments, the trust index determining server 107 may be configured in a cloud environment.
Each of the plurality of computing devices 103 comprises a processor 109, an Input/output (I/O) interface 111 and a memory 113 as shown in the FIG.1B. For the ease of understanding, the present disclosure is explained further for a single computing device. The same method is followed by each of the plurality of computing devices 103. In an embodiment, the I/O interface 111 of the computing device 103 may receive a data request related to a subscriber, from the trust index determining server 107. In some embodiments, the subscriber may include, but not limited to, a customer, a service provider or a Small/Medium Enterprise (SME). Further, the processor 109 of the computing device 103 may retrieve contextual data related to the subscriber from the corresponding data repository 101. In some embodiments, the contextual data may include, but not limited to, sensitive contextual data and non-sensitive contextual data. In some embodiments, the contextual data may be retrieved upon receiving a consent from the subscriber to access data records related to the subscriber. In some embodiments, the consent of the subscriber may be implicit in nature, in order to avail a digital identity trust index. Further, the processor 109 of the computing device 103 may generate secure sensitive contextual data using a noise perturbed feature based technique on the retrieved sensitive contextual data. Thereafter, the processor 109 of the computing device 103 may provide the secure sensitive contextual data and the non-sensitive contextual data to the corresponding third party analytics agent 105, to generate one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data and the non-sensitive contextual data. The processor 109 of the computing device 103 may then transmit the one or more intermediate analytics results to the trust index determining server 107 for determining a digital identity trust index.
The trust index determining server 107 comprises a processor 115, an Input/output (I/O) interface 117 and a memory 119 as shown in FIG.1C. In an embodiment, the I/O interface 117 of the trust index determining server 107 may receive one or more intermediate analytics results from the plurality of computing devices 103 associated with the corresponding plurality of data repositories 101. In some embodiments, the one or more intermediate results may include, but not limited to, one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features. The processor 115 of the trust index determining server 107 may determine the digital identity trust index of the subscriber by providing the one or more intermediate results to a machine learning model of the trust index determining server 107. The digital identity trust index thus determined may be stored in the memory 119 of the trust index determining server 107.
FIG.2A shows a detailed block diagram of a computing device for generating analytics results for determining a digital identity trust index in digital ecosystem in accordance with some embodiments of the present disclosure.
In some implementations, the computing device 103 may include data 203 and modules 205. As an example, the data 203 of the computing device 103 are stored in the memory 113 configured in the computing device 103 as shown in the FIG.2A. In one embodiment, the data 203 of the computing device 103 may include contextual data 207, secure sensitive contextual data 209, intermediate result data 211 and other data 213. In the illustrated FIG.2A, modules 205 of the computing device 103 are described herein in detail.
In some embodiments, the data 203 of the computing device 103 may be stored in the memory 113 of the computing device 103 in the form of various data structures. Additionally, the data 203 of the computing device 103 can be organized using data models, such as relational or hierarchical data models. In some embodiments, the other data 213 may store data, including temporary data and temporary files, generated by the modules 205 of the computing device 103 for performing the various functions of the computing device 103.
In some embodiments, the data 203 of the computing device 103 stored in the memory 113 may be processed by the modules 205 of the computing device 103. The modules 205 may be stored within the memory 113 of the computing device 103. In an example, the modules 205 communicatively coupled to the processor 109 configured in the computing device 103, may also be present outside the memory 113 as shown in FIG.2A and implemented as hardware. As used herein, the term modules refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or group) and memory that execute one or more software or firmware programs, a combinational logic circuit, and/or other suitable components that provide the described functionality.
In an embodiment, the modules 205 of the computing device 103 may include, for example, a receiving module 221, data retrieving module 223, a secure data generating module 225, an intermediate result generating module 227, a transmitting module 229 and other modules 231. The other modules 231 may be used to perform various miscellaneous functionalities of the computing device 103. It will be appreciated that such aforementioned modules 205 of the computing device 103 may be represented as a single module or a combination of different modules.
In some embodiments, the receiving module 221 may receive a data request from a trust index determining server 107 associated with the computing device 103. The data request may be related to a subscriber. In some embodiments, the receiving module 221 may receive the data request from the trust index determining server 107, when the subscriber wishes to avail a service related to a business application. In some embodiments, the trust index determining server 107 may send the data request to some of the plurality of computing devices 103 i.e. to some of the plurality of data repositories based on a kind of service which the subscriber wishes to avail. In some other embodiments, the trust index determining server 107 may send the data request to each of the plurality of computing devices 103, irrespective of the kind of service that the subscriber wishes to avail. In yet other embodiments, the trust index determining server 107 may send the data request to some of the plurality of computing devices 103 i.e. to some of the plurality of data repositories based on the kind of subscriber. For example, when the subscriber is a customer, the data request may be sent to tax records repository, as the tax records are relevant for the customer. However, when the subscriber is a service provider or a Small/Medium Entity (SME), the data request may be sent to GST records repository since the GST records are relevant for the service providers and the SMEs, and not the tax records.
In some embodiments, the data retrieving module 223 may retrieve the contextual data 207 related to the subscriber, from a corresponding data repository 101 of the plurality of data repositories 101. In some embodiments, the contextual data may include, but not limited to, at least one of sensitive contextual data and non-sensitive contextual data. Sensitive contextual data may be the contextual data 207 related to the subscriber, which is confidential in nature. As an example, in the context of financial transactions, data related to date and time of financial transactions, transaction amount, and Automated Teller Machine (ATM)/Point Of Sale (POS)/bank location of the transaction and the like, may be regarded as sensitive contextual data of the subscriber. Non-sensitive contextual data may be the contextual data 207 related to the subscriber, which is non-confidential in nature, and easily available in public. As an example, name, age, gender, educational qualification and the like, may be regarded as non-confidential data of the subscriber.
In some embodiments, the secure data generating module 225 may generate secure sensitive contextual data 209 by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique. As an example, if the sensitive contextual data is related to financial transactions, the contextual elements of the sensitive contextual data may be “date and time of the transactions”, “transaction amount”, and “location of the transaction” and the like. As an example, if the sensitive contextual data is related to health records, the contextual elements of the sensitive contextual data may be “hospital location”, “disease diagnosed”, and “date and time of doctor visit” and the like.
In some embodiments, to perform the noise perturbed feature based technique based on the degree of sensitivity of each contextual element, the secure data generating module 225 may divide at least one of the sensitive contextual data and the non-sensitive contextual data into plurality of matrices. In some embodiments, the secure data generating module 225 may divide based on total number of the contextual elements of at least one of the sensitive contextual data and the non-sensitive contextual data. In some embodiments, contextual elements of the sensitive contextual data may be referred as sensitive contextual elements. Further, the secure data generating module 225 may cluster a plurality of sensitive contextual elements among total number of sensitive contextual elements, into a plurality of clusters, based on a predefined clustering technique. As an example, the predefined clustering technique may include, but not limited to, K-means clustering technique. Subsequently, the secure data generating module 225 may compute a centroid value for each of the plurality of clusters. Further, the secure data generating module 225 may replace a value of, each of the plurality of sensitive contextual elements with a combined value obtained by combining the centroid value of each of the corresponding plurality of clusters and a first corresponding random noise. In some embodiments, the secure data generating module 225 may select the first random noise for each of the plurality of sensitive contextual elements from a Laplace distribution, based on the degree of sensitivity of each of the plurality of sensitive contextual elements. In some embodiments, the degree of sensitivity of each contextual element may be quantified into one of “Critical”, “High”, “Medium” and “Low” (also referred as categories of degree of sensitivity). In some embodiments, “Critical” may be the highest order of degree of sensitivity among four categories of the degree of sensitivity mentioned above, followed by the order “High”, “Medium” and “Low”. “Low” is the lowest order of degree of sensitivity among the four categories of the degree of sensitivity mentioned above. In some embodiments, the scale of noise selected for a sensitive contextual element from the Laplace distribution is directly proportional to the degree of sensitivity of the sensitive contextual element. For example, if the degree of sensitivity of the sensitive contextual element is “Critical”, then noise of highest scale is selected from the Laplace distribution. Similarly, if the degree of sensitivity of the sensitive contextual element is “Low”, then noise of lowest scale is selected from the Laplace distribution. In some embodiments, the secure data generating module 225 may determine the degree of sensitivity of each contextual element by initially determining a predefined category of each of the contextual element. Each predefined category may be associated with a degree of sensitivity. Thereafter, the secure data generating module 225 may assign each contextual element of the sensitive contextual data, with the degree of sensitivity of the corresponding predefined category.
Further, the secure data generating module 225 may replace a value of each of rest of the plurality of sensitive contextual elements among the total number of sensitive contextual elements with a combined value obtained by combining original value of each corresponding sensitive contextual element and a second corresponding random noise. In some embodiments, the secure data generating module 225 may select the second random noise for each of rest of the plurality of sensitive contextual elements from the Laplace distribution, based on the degree of sensitivity of each of the rest of the plurality of sensitive contextual elements. The data thus generated may be stored as secure sensitive contextual data 209.
As an example, consider the following sensitive contextual elements:
SC: Sensitive-Critical
SH: Sensitive-High
SM: Sensitive-Medium
SL: Sensitive-Low
Further, consider non-sensitive contextual elements that are represented as NSF.
Consider, total number of sensitive contextual elements: N=4 and total number of non-sensitive contextual elements: M=3
In some embodiments, the secure data generating module 225 divides the sensitive contextual data and the non-sensitive contextual data into a plurality of matrices, wherein the secure data generating module 225 may determine size of each matrix using the below Equation 1.
(N + M) x (N + M)2 ------------- Equation 1
In the above Equation 1,
N represents number of sensitive contextual elements in the sensitive contextual data;
M represents number of non-sensitive contextual elements in the non-sensitive contextual data.
In the above example, N=4 and M=3. Therefore, using the Equation 1, the secure data generating module 225 divides the sensitive contextual data and the non-sensitive contextual data into plurality of matrices, each matrix of size:
7 (Columns) x 49 (Rows)
The exemplary matrix with 7 columns and 49 rows is as shown in FIG.2B.
In the exemplary matrix shown above, each row denotes a user record and each column denotes the sensitive contextual element/non-sensitive contextual element.
Further, the secure data generating module 225 may select plurality of sensitive contextual elements using the below Equation 2.
K = floor(sqrt(N)) -------------- Equation 2
In the above Equation 2,
N denotes the total number of sensitive contextual elements; and
K denotes number of columns of the matrix comprising sensitive contextual elements to be selected.
In the above example, since total number of sensitive contextual elements = 4, using the Equation 2, K i.e. number of columns of sensitive contextual elements selected would be 2.
Thereafter, the secure data generating module 225 may randomly select K number of columns comprising sensitive contextual elements from the total number of columns comprising the sensitive contextual elements in the matrix.
In the above example, consider, the secure data generating module 225 selects columns 1 and 3 of the matrix. Upon using a predefined clustering technique, consider the secure data generating module 225 clusters the values of the sensitive contextual elements present in the selected columns 1 and 3 of the matrix into plurality of clusters, and computes centroid value for each of the plurality of clusters, as shown below:
- Column 1: Based on values of sensitive contextual elements:
Divided into X = 4 Clusters
Centroids of each Cluster - C1,1, C1,2, C1,3, C1,4.
- Column 3: Based on values of sensitive contextual elements:
Divided into X = 4 Clusters
Centroids of each Cluster - C3,1, C3,2, C3,3, C3,4.
Size of Cluster: No. of Rows in Cluster: Consider, Cluster 1 = 12, Cluster 2 = 12, Cluster 3 = 12, Cluster 4 = 13.
The exemplary four clusters formed in the above example, are as shown in the FIG.2C.
Thereafter, the secure data generating module 225 may replace value of each of the plurality of sensitive contextual elements with a combined value obtained by combining the centroid value of each of the corresponding plurality of clusters and a first corresponding random noise. As an example, consider, original value in Column 1: Row 1 is V1,1 as shown in the FIG.2B. The original value V1,1 may be thus replaced by the combined value (C1,1 + N1,1), in which C1,1 denotes centroid value of the corresponding cluster and N1,1 is the first corresponding random noise selected from a Laplace distribution. The first random noise may be selected from the Laplace distribution based on the degree of sensitivity. For instance, if the degree of sensitivity of the sensitive contextual element is “Critical”, then noise of highest scale is selected from the Laplace distribution. Similarly, the secure data generating module 225 replaces each value of the sensitive contextual elements in the columns 1 and 3 of the matrix. The exemplary values that are replaced in columns 1 and 3 of the matrix are as shown in the FIG.2C.
Further, for the remaining columns comprising sensitive contextual elements in the matrix, apart from columns 1 and 3, the secure data generating module 225 may replace each value of the sensitive contextual elements of these remaining columns of the matrix, with a combined value obtained by combining original value of each corresponding sensitive contextual element and a second random noise selected from the Laplace distribution. As an example, the remaining columns comprising sensitive contextual elements in the matrix may be columns 2 and 4. As an example, the original value of the column 2: Row 2 is V2,2, as shown in the FIG.2B. The original value V2,2 may be thus replaced by the value (V2,2 + N2,2), which means that second random noise N2,2 is added to the original value V2,2. Similarly, the secure data generating module 225 replaces each value of the sensitive contextual elements in the remaining columns 2 and 4 of the matrix. The exemplary values that are replaced in columns 2 and 4 of the matrix are as shown in the FIG.2C.
Further, the intermediate result generating module 227 may generate one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data 209 and the non-sensitive contextual data. The intermediate result generating module 227 may generate the federated analytics using a third party analytics agent 105 associated with the computing device 103. In some embodiments, the intermediate result generating module 227 may generate the one or more intermediate analytics results by initially identifying one or more candidate features required for determining the digital identity trust index of the subscriber, and one or more sub-features related to each of the one or more candidate features, from at least one of the secure sensitive contextual data 209 and the non-sensitive contextual data. In some embodiments, the one or more candidate features and the one or more sub-features related to each of the one or more candidate features may be determined based on an optimal feature set. In some embodiments, optimal feature set may be a subset of features of the secure sensitive contextual data 209, wherein the optimal feature set may include the most relevant and important features required for determining the digital identity trust index. In some embodiments, the optimal feature set is determined during a training phase of a machine learning model used in the trust index determining server 107. The optimal feature set is thereafter configured in each of the plurality of third party analytics agents 105 for performing the federated analytics. An exemplary optimal feature set for contextual data retrieved from each of the plurality of repositories 101 is as shown in the below Table 1.
Sl.no Data Repository Optimal feature set
1. Financial transactions Repository 1. Credit payment history (late payments, payment defaults)
2. Current debt Burden
3. Business market value assessment
4. Business presence in multiple sectors
5. Financial transactions from multiple banks
2. Tax/Goods and Services Tax (GST) records Repository 1. GST payment defaults
2. Property tax
3. Tax bracket
3. Insurance data repository 1. Amount insured
2. Past insurance claims
3. Types of insurance (Financial/Health/Property)
4. Insurance payment defaults
4. Investment records repository 1. Investment portfolio
2. Premium payment defaults
5. Social Media data repository 1. Behavioural analysis
6. Demographic statistics repository 1. Business demand supply information
2. Employment/Financial status
3. Infrastructure Quality
4. Population density
7. Identity verification repository 1. National Identity verification
2. Business registration
3. E-verified documents
8. Health stack information repository 1. Customer health record categorization
2. Health service provider records
9. Police, Central Bureau of Investigation (CBI) and judiciary records repository 1. Criminal background status
10. Online customer feedback repository 1. Customer ratings
11. Service quality records repository 1. Service provider quality of services
12. Blockchain records repository 1. Land/Property records
2. Education certificates
3. E-driving license and Records
4. E-passport and Travel records
13. Telecom data repository 1. Location based service records
Table 1
Further, the intermediate result generating module 227 may determine a normalized value with respect to a predefined scale, for each of the one or more candidate features using the one or more sub-features related to each of the one or more candidate features. In some embodiments, the one or more candidate features and the normalized value corresponding to each of the one or more candidate features constitute the one or more intermediate results. The one or more intermediate results are stored as the intermediate result data 211.
As an example, if “financial transactions” is considered as the candidate feature 1 for determining the digital trust index, the one or more sub-features of the candidate feature “financial transaction” may be
? Date-Time: [Input to Candidate feature 1 Value]
? ATM/POS/Bank Location: [Input to Candidate feature 1 Value]
? Transaction Amount: [Input to Candidate feature 1 Value]
Similarly, if “credit payment history and default” is considered as the candidate feature 2 for determining the digital trust index, the one or more sub-features of the candidate feature “credit payment history and default” may be
? Date-Time: [Input to Candidate feature 2 Value]
? Credit Amount: [Input to Candidate feature 2 Value]
? Annual Late Payments: [Input to Candidate feature 2 Value]
? Annual Payment Defaults [Input to Candidate feature 2 Value]
For each of the one or more exemplary candidate features, the intermediate result generating module 227 may determine a normalized value with respect to a predefined scale using the one or more exemplary sub-features related to each of the one or more exemplary candidate features. As an example, the predefined scale may be between 0-100. In some embodiments, normalization is performed in order to convert general numeric values of the one or more candidate features into numeric values corresponding to the predefined scale. In some embodiments, machine learning models work more efficiently with normalized values than general numeric values.
Further, in some embodiments, the transmitting module 229 may transmit each of the one or more intermediate analytics results to the trust index determining server 107 associated with the computing device 103. In some embodiments, till now the present disclosure has been described in view of one computing device 103, one data repository associated with the computing device 103 and one third party analytics agent associated with the computing device 103. The same procedure is followed for each of the plurality of computing devices 103 and the corresponding plurality of data repositories 101 and the corresponding plurality of third party analytics agents 105.
Therefore, each of the plurality of computing devices 103 that received the data request from the trust index determining server 107, transmit each of the one or more intermediate analytics results to the trust index determining server 107, for determining a digital identity trust index in digital ecosystems.
FIG.2D shows a detailed block diagram of a trust index determining server for determining a digital identity trust index in digital ecosystems in accordance with some embodiments of the present disclosure.
In some implementations, the trust index determining server 107 may include data 235 and modules 237. As an example, the data 235 of the trust index determining server 107 is stored in the memory 119 configured in the trust index determining server 107 as shown in the FIG.2D. In one embodiment, the data 235 of the trust index determining server 107 may include intermediate result data 211, predicted data 239 and other data 241. In the illustrated FIG.2D, modules 237 of the trust index determining server 107 are described herein in detail.
In some embodiments, the data 235 may be stored in the memory 119 in form of various data structures. Additionally, the data 235 can be organized using data models, such as relational or hierarchical data models.
In some embodiments, the data 235 of the trust index determining server 107 may be stored in the memory 119 of the trust index determining server 107 in the form of various data structures. Additionally, the data 235 of the trust index determining server 107 can be organized using data models, such as relational or hierarchical data models. In some embodiments, the other data 235 of the trust index determining server 107 may store data, including temporary data and temporary files, generated by the modules 237 of the trust index determining server 107 for performing the various functions of the trust index determining server 107.
In some embodiments, the data 235 of the trust index determining server 107 stored in the memory 119 may be processed by the modules 237 of the trust index determining server 107. The modules 237 of the trust index determining server 107 may be stored within the memory 119 of the trust index determining server 107. In an example, the modules 237 communicatively coupled to the processor 115 configured in the trust index determining server 107, may also be present outside the memory 119 as shown in FIG.2D and implemented as hardware. As used herein, the term modules refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or group) and memory that execute one or more software or firmware programs, a combinational logic circuit, and/or other suitable components that provide the described functionality.
In an embodiment, the modules 237 of the trust index determining server 107 may include, for example, a receiving module 243, a feature selection module 245, a trust index determining module 247, a trust index updating module 249 and other modules 251. The other modules 251 may be used to perform various miscellaneous functionalities of the trust index determining server 107. It will be appreciated that such aforementioned modules 237 of the trust index determining server 107 may be represented as a single module or a combination of different modules.
In some embodiments, the receiving module 243 may receive one or more intermediate analytics results from a plurality of computing devices 103 associated with a corresponding plurality of data repositories 101. In some embodiments, the one or more intermediate results are related to a subscriber whose digital identity trust index is to be determined. The one or more intermediate results may include, but not limited to, one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features.
Further, the feature selection module 245 may select a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set (This phase may be referred as prediction phase). In some embodiments, the key feature of the one or more candidate features, and the rest of the one or more candidate features are determined during a training phase of the machine learning model, based on historical data related to an application for which the digital identity trust index of the subscriber is being determined. In some embodiments, the historical data used for training the machine learning model is data collected for a large number of users. In some embodiments, supervised machine learning models may be used in the present disclosure.
As an example, consider digital identity trust index is being determined for financial credit assessment application. During the training phase, the machine learning model may be trained using historical data features such as “credit payment history” as an output tag and rest of the features as an input tag, to predict the probability of credit payment, which would be considered as the digital identity trust index of the user for financial credit assessment application. As an example, the rest of the features for the financial credit assessment application may be “financial transactions from multiple banks”, “current debt burden”, “business market value assessment” and the like.
Therefore, during the training phase, the machine learning model is trained by tagging the main feature whose prediction is required, with an output tag and rest of the features with an input tag.
Referring back to the prediction phase, the trust index determining module 247 may generate a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features, as an input, to the machine learning model. In some embodiments, based on the training that the machine learning model has undergone during the training phase, the machine learning model may generate the prediction result of the key feature tagged as output, based on the optimal feature set. In some embodiments, the generated prediction result is a probability associated with the key feature tagged as “Output”, which is considered as the digital identity trust index of the subscriber. The prediction result generated by the trust index determining module 247 may be stored as the predicted data 239. As an example, in a business assessment application, “business continuity” may be selected as the key feature among the one or more candidate features, and therefore is tagged as “Output”. However, rest of the candidate features such as “business presence”, “business market value”, “business debts” and the like may be tagged as “Input”. Upon providing the rest of the candidate features to the machine learning model during the prediction phase, the machine learning model may predict the probability of business continuity, which is the probability of the key feature tagged as “Output”. The prediction result i.e. probability of business continuity may be considered as digital identity trust index of the subscriber for the application of business assessment.
Further, in some embodiments, the trust index updating module 249 may determine an updated digital identity trust index, periodically to reflect changes in the information related to the subscriber. The periodicity may be pre-set by a service provider of the digital identity trust index. In some embodiments, the updated digital identity trust index may be determined by determining an exponential moving average of currently determined digital identity trust index based on updated contextual data, and the previously determined digital identity trust index.
The current digital identity trust score may be determined using the below Equation 3.
Current_Digital_Identity_Trust_Index =
Digital_Identity_Trust_Index [n] * Exp_Percent + Digital_Identity_Trust_Index [n-1] * (1 - Exp_Percent) ------ Equation 3
In the above Equation 3,
Current_Digital_Identity_Trust_Index denotes the current value of digital identity trust index after applying Exponential Moving Average.
Digital_Identity_Trust_Index [n] denotes nth computation of the digital identity trust index, where n= 1, 2, 3 ….
Digital_Identity_Trust_Index [n-1] denotes n-1th computation of the digital identity trust index, where n-1 = 1, 2, 3….
Time_Period denotes time duration between consecutive computations of the digital identity trust index, or in other words, periodicity set by the service provider for updating the digital identity trust index.
Exp_Percent denotes a decay constant.
In some embodiments, the trust index updating module 249 may determine the value of Exp_Percent based on the below Equation 4.
Exp_Percent = 2 / (Time_Period + 1) ------------- Equation 4
Henceforth, the process of determining a digital identity trust index in a digital ecosystem is explained with the help of one or more examples for better understanding of the present disclosure. However, the one or more examples should not be considered as limitation of the present disclosure.
Consider an exemplary application such as financial credit assessment application. In some embodiments, financial credit assessment application is just exemplary and should not be construed as a limitation of the present disclosure. The method disclosed in the present disclosure can be applicable for wide variety of applications such as employee on-boarding trust score, mobility service provider trust score, demographic trust score for smart city planning, data trust score for information sharing, logistics service provider trust score and the like.
Consider the subscriber for whom the digital identity trust index has to be determined is a Small/Medium Entity (SME). Consider the subscriber 253 approaches a financial entity 255 for a loan. The financial entity 255 may avail the services of the trust index determining server 107 to determine the digital identity trust index of the subscriber 253 in real-time, as shown in the FIG.2E.
Thereafter, the trust index determining server 107 may send a data request to exemplary data repositories such as financial transactions repository 1011, identity verification repository 1012 and police records repository 1013. Consider the computing devices 1031, 1032, and 1033 are associated with the financial transactions repository 1011, identity verification repository 1012 and police records repository 1013, respectively, as shown in the FIG.2F. Upon receiving the data request, each of the computing devices 1031, 1032, and 1033 may retrieve contextual data from the corresponding data repositories 1011, 1012, and 1013.
As an example, consider the contextual data retrieved from the financial transactions repository 1011 comprises the following data of the subscriber 253 related to the financial transactions, as shown in the below Table 2. Also, the below Table 2 shows whether the retrieved contextual data is sensitive/non-sensitive, along with its degree of sensitivity.
Sl.no Contextual data (also referred as candidate features) Contextual elements (also referred as sub-features of the candidate features) Sensitive Non-sensitive Degree of sensitivity
1. Financial transactions from multiple banks 1. Date and time
Yes - Critical
2. ATM/POS/Bank location
Yes - Critical
3. Transaction Amount Yes - High
2. Credit payment history and defaults 1. Date and time
Yes - Medium
2. Credit Amount - Yes -
3. Annual late payments
- Yes -
4. Annual payment defaults - Yes -
3. Current Debt Burden 1. No. of entities from which debt is already taken
- Yes -
2. Debt burden amount - Yes -
4. Business market value assessment Assessment amount - Yes -
5. Business presence in multiple sectors Number of sectors - Yes -
Table 2
Similarly, the contextual data may be retrieved from identity verification repository 1012 and police records repository 1013. Upon retrieving the contextual data, the computing devices 1031, 1032, and 1033 may generate secure sensitive contextual data 209 by applying noise perturbed feature based technique on the sensitive contextual data retrieved from the corresponding data repositories 1011, 1012, and 1013. Thereafter, the computing devices 1031, 1032, and 1033 may provide the corresponding non-sensitive contextual data and the secure sensitive contextual data 209 to the corresponding third party analytics agents 1051, 1052, and 1053, respectively, as shown in the FIG.2F. Each of the third party analytics agents 1051, 1052, and 1053 may generate one or more intermediate analytics results based on the one or more candidate features and the one or more sub-features related to each of the one or more candidate features, that are determined based on an optimal feature set. The one or more intermediate analytics results of the one or more candidate features and the one or more sub-features related to each of the one or more candidate features shown in the above Table 2, are given below:
o Financial Transactions from Multiple Banks: [Candidate Feature 1 (CF 1) Value: On Scale of 100]
? Date-Time: [Input to CF 1 Value]
? ATM/POS/Bank Location: [Input to CF 1 Value]
? Transaction Amount: [Input to CF 1 Value]
o Credit Payment History and Defaults: [CF 2 Value: On Scale of 100]
? Date-Time: [Input to CF 2 Value]
? Credit Amount: [Input to CF 2 Value]
? Annual Late Payments: [Input to CF 2 Value]
? Annual Payment Defaults [Input to CF 2 Value]
o Current Debt Burden: [CF 3 Value: On Scale of 100]
? Number of Entities from which Credit is already taken: [Input to CF 3 Value]
? Debt Burden Amount: [Input to CF 3 Value]
o Business Market Value Assessment: [CF 4 Value: On Scale of 100]
? Assessment Amount: [Input to CF 4 Value]
o Business Presence in Multiple Sectors: [CF 5 Value: On Scale of 100]
? Number of Sectors (Finance/Travel/Ecommerce): [Input to CF 5 Value]
Similarly, the one or more intermediate analytics results may be determined for the one or more candidate features and the one or more sub-features related to the one or more candidate features of the contextual data retrieved from identity verification repository 1012 and police records repository 1013.
The computing devices 1031, 1032, and 1033 may transmit the corresponding one or more intermediate analytics results to the trust index determining server 107. Among the one or more intermediate analytics results, the trust index determining server 107 may select a key feature among the one or more candidate features, based on the optimal feature set.
Since, the application is a financial credit assessment application, the trust index determining server 107 selects “Credit payment history and defaults” as the key feature among the one or more candidate features, based on the optimal feature set. Therefore, the key feature and rest of the one or more candidate features are tagged as shown in the below Table 3.
Candidate feature type Candidate feature Tag assigned to the candidate feature
Key feature Credit payment history and defaults Output
Rest of the one or more candidate features Financial transactions from multiple banks
Input
Current Debt Burden
Input
Business market value assessment
Input
Business presence in multiple sectors
.
.
. Input
Table 3
Further, the trust index determining server 107 may predict a probability associated with the key feature tagged as “Output”, i.e. probability of the credit payment, by providing the rest of the one or more candidate features tagged as “Input” to the machine learning model of the trust index determining server 107. The prediction result i.e. the numeric value thus generated for probability of the credit payment, is considered as the digital identity trust index of the subscriber for the financial credit assessment application.
FIG.3A shows flowchart illustrating a method of generating analytics results for determining a digital identity trust index in digital ecosystem, in accordance with some embodiments of the present disclosure.
As illustrated in FIG.3A, the method 300a includes one or more blocks illustrating a method of generating analytics results for determining a digital identity trust index in digital ecosystem. The method 300a may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform functions or implement abstract data types.
The order in which the method 300a is described is not intended to be construed as a limitation, and any number of the described method blocks can be combined in any order to implement the method 300a. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described herein. Furthermore, the method 300a can be implemented in any suitable hardware, software, firmware, or combination thereof.
At block 301, the method includes receiving, by a processor 109 of the computing device 103, a data request from a trust index determining server 107. In some embodiments, the data request is related to a subscriber such as a customer, a service provider or a small/medium enterprise.
At block 303, the method includes retrieving, by the processor 109 of the computing device 103, contextual data related to the subscriber, from a corresponding data repository of the plurality of data repositories 101. In some embodiments, the contextual data may include at least one of sensitive contextual data and non-sensitive contextual data.
At block 305, the method includes generating, by the processor 109 of the computing device 103, secure sensitive contextual data 209 by introducing random noise to each contextual element of the sensitive contextual data, based on a degree of sensitivity of each contextual element in the sensitive contextual data, using a noise perturbed feature based technique. In some embodiments, the degree of sensitivity of each contextual element may be quantified into one of “Critical”, “High”, “Medium” and “Low”.
At block 307, the method includes generating, by the processor 109 of the computing device 103, one or more intermediate analytics results by performing federated analytics on at least one of the secure sensitive contextual data 209 and the non-sensitive contextual data, using a third party analytics agent 105 associated with the computing device 103. In some embodiments, the processor 109 of the computing device 103 may transmit the one or more intermediate analytics results to the trust index determining server 107, for determining a digital identity trust index in digital ecosystems.
FIG.3B shows flowchart illustrating a method of determining a digital identity trust index in digital ecosystems, in accordance with some embodiments of the present disclosure.
As illustrated in FIG.3B, the method 300b includes one or more blocks illustrating a method of determining a digital identity trust index in digital ecosystems. The method 300b may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform functions or implement abstract data types.
The order in which the method 300b is described is not intended to be construed as a limitation, and any number of the described method blocks can be combined in any order to implement the method 300b. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described herein. Furthermore, the method 300b can be implemented in any suitable hardware, software, firmware, or combination thereof.
At block 311, the method includes receiving, by a processor 115 of a trust index determining server 107, one or more intermediate analytics results from a plurality of computing devices 103 associated with a corresponding plurality of data repositories 101. In some embodiments, the one or more intermediate results are related to a subscriber whose digital identity trust index is to be determined. In some embodiments, the one or more intermediate results may include, one or more candidate features required for determining the digital identity trust index of the subscriber and a normalized value corresponding to each of the one or more candidate features.
At block 313, the method includes selecting, by the processor 115 of the trust index determining server 107, a key feature of the one or more candidate features as “Output” and rest of the one or more candidate features as “Input”, based on an optimal feature set. In some embodiments, the optimal feature set, the key feature of the one or more candidate features, and the rest of the one or more candidate features may be determined during a training phase of the machine learning model, based on historical data related to an application for which the digital identity trust index of the subscriber is being determined.
At block 315, the method includes generating, by the processor 115 of the trust index determining server 107, a prediction result of the key feature tagged as “Output” by providing the rest of the one or more candidate features tagged as “Input” and the normalized value corresponding to each of the rest of the one or more candidate features as an input to a machine learning model, for determining the digital identity trust index of the subscriber. In some embodiments, the machine learning model may be trained to generate the prediction result based on the optimal feature set. In some embodiments, the processor 115 of the trust index determining server 107 may determine an updated digital identity trust index, periodically, by determining an exponential moving average of currently determined digital identity trust index based on updated contextual data, and the previously determined digital identity trust index.
FIG.4 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.
In some embodiments, FIG.4 illustrates a block diagram of an exemplary computer system 400 for implementing embodiments consistent with the present invention. In some embodiments, the computer system 400 can be a computing device 103 used for generating analytics results for determining a digital identity trust index in digital ecosystem. In some embodiments, the computer system 400 can be a trust index determining server 107 used for determining a digital identity trust index in digital ecosystems. FIG.4 is illustrated considering the computer system 400 as the computing device 103. However, this should not be considered as a limitation, since the computer system 400 can be the trust index determining server 107 as well (not shown in FIG.4). The computer system 400 may include a central processing unit (“CPU” or “processor”) 402. The processor 402 may include at least one data processor for executing program components for executing user or system-generated business processes. A user may include a person, a person using a device such as those included in this invention, or such a device itself. The processor 402 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc.
The processor 402 may be disposed in communication with input devices 411 and output devices 412 via I/O interface 401. The I/O interface 401 may employ communication protocols/methods such as, without limitation, audio, analog, digital, stereo, IEEE-1394, serial bus, Universal Serial Bus (USB), infrared, PS/2, BNC, coaxial, component, composite, Digital Visual Interface (DVI), high-definition multimedia interface (HDMI), Radio Frequency (RF) antennas, S-Video, Video Graphics Array (VGA), IEEE 802.n /b/g/n/x, Bluetooth, cellular (e.g., Code-Division Multiple Access (CDMA), High-Speed Packet Access (HSPA+), Global System For Mobile Communications (GSM), Long-Term Evolution (LTE), WiMax, or the like), etc.
Using the I/O interface 401, computer system 400 may communicate with input devices 411 and output devices 412.
In some embodiments, the processor 402 may be disposed in communication with a communication network 409 via a network interface 403. The network interface 403 may communicate with the communication network 409. The network interface 403 may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10/100/1000 Base T), Transmission Control Protocol/Internet Protocol (TCP/IP), token ring, IEEE 802.11a/b/g/n/x, etc. Using the network interface 403 and the communication network 409, the computer system 400 may communicate with plurality of data repositories 101 (1011 up to 101n), plurality of third party analytics agents 105 (1051 up to 105n) and a trust index determining server 107. In some embodiments, the computer system 400 may communicate via IP based communication and non-internet or non-IP based communication such as Universal Serial Bus (USB), Bluetooth and the like. The communication network 409 can be implemented as one of the different types of networks, such as intranet or Local Area Network (LAN), Closed Area Network (CAN) and such within the autonomous vehicle. The communication network 409 may either be a dedicated network or a shared network, which represents an association of the different types of networks that use a variety of protocols, for example, Hypertext Transfer Protocol (HTTP), CAN Protocol, Transmission Control Protocol/Internet Protocol (TCP/IP), Wireless Application Protocol (WAP), etc., to communicate with each other. Further, the communication network 409 may include a variety of network devices, including routers, bridges, servers, computing devices, storage devices, etc. In some embodiments, the processor 402 may be disposed in communication with a memory 405 (e.g., RAM, ROM, etc. not shown in FIG.4) via a storage interface 404. The storage interface 404 may connect to memory 405 including, without limitation, memory drives, removable disc drives, etc., employing connection protocols such as Serial Advanced Technology Attachment (SATA), Integrated Drive Electronics (IDE), IEEE-1394, Universal Serial Bus (USB), fibre channel, Small Computer Systems Interface (SCSI), etc. The memory drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory devices, solid-state drives, etc.
The memory 405 may store a collection of program or database components, including, without limitation, a user interface 406, an operating system 407, a web browser 408 etc. In some embodiments, the computer system 400 may store user/application data, such as the data, variables, records, etc. as described in this invention. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle or Sybase.
The operating system 407 may facilitate resource management and operation of the computer system 400. Examples of operating systems 407 include, without limitation, APPLE® MACINTOSH® OS X®, UNIX®, UNIX-like system distributions (E.G., BERKELEY SOFTWARE DISTRIBUTION® (BSD), FREEBSD®, NETBSD®, OPENBSD, etc.), LINUX® DISTRIBUTIONS (E.G., RED HAT®, UBUNTU®, KUBUNTU®, etc.), IBM®OS/2®, MICROSOFT® WINDOWS® (XP®, VISTA®/7/8, 10 etc.), APPLE® IOS®, GOOGLETM ANDROIDTM, BLACKBERRY® OS, or the like. The User interface 406 may facilitate display, execution, interaction, manipulation, or operation of program components through textual or graphical facilities. For example, user interfaces 406 may provide computer interaction interface elements on a display system operatively connected to the computer system 400, such as cursors, icons, checkboxes, menus, scrollers, windows, widgets, etc. Graphical User Interfaces (GUIs) may be employed, including, without limitation, Apple® Macintosh® operating systems’ Aqua®, IBM® OS/2®, Microsoft® Windows® (e.g., Aero, Metro, etc.), web interface libraries (e.g., ActiveX®, Java®, Javascript®, AJAX, HTML, Adobe® Flash®, etc.), or the like.
In some embodiments, the computer system 400 may implement the web browser 408 stored program components. The web browser 408 may be a hypertext viewing application, such as MICROSOFT® INTERNET EXPLORER®, GOOGLETM CHROMETM, MOZILLA® FIREFOX®, APPLE® SAFARI®, etc. Secure web browsing may be provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. Web browsers 408 may utilize facilities such as AJAX, DHTML, ADOBE® FLASH®, JAVASCRIPT®, JAVA®, Application Programming Interfaces (APIs), etc. In some embodiments, the computer system 400 may implement a mail server stored program component. The mail server may be an Internet mail server such as Microsoft Exchange, or the like. The mail server may utilize facilities such as Active Server Pages (ASP), ACTIVEX®, ANSI® C++/C#, MICROSOFT®, .NET, CGI SCRIPTS, JAVA®, JAVASCRIPT®, PERL®, PHP, PYTHON®, WEBOBJECTS®, etc. The mail server may utilize communication protocols such as Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), MICROSOFT® exchange, Post Office Protocol (POP), Simple Mail Transfer Protocol (SMTP), or the like. In some embodiments, the computer system 400 may implement a mail client stored program component. The mail client may be a mail viewing application, such as APPLE® MAIL, MICROSOFT® ENTOURAGE®, MICROSOFT® OUTLOOK®, MOZILLA® THUNDERBIRD®, etc.
Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present invention. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., non-transitory. Examples include Random Access Memory (RAM), Read-Only Memory (ROM), volatile memory, non-volatile memory, hard drives, Compact Disc (CD) ROMs, Digital Video Disc (DVDs), flash drives, disks, and any other known physical storage media.
Advantages of the present disclosure are illustrated herein.
The present disclosure associates a digital identity trust index with subscribers such as customers, service providers and SMEs based on contextual interpretation of data records retrieved from wide variety of data repositories, thereby providing the advantages of:
• Easy interpretation of digital trust level of the subscriber.
• Increased scope of verification.
• Accurate comprehension of digital trust index.
• Reduction in fraud cases.
• Early risk management.
The present disclosure secures the sensitive contextual data using a noise perturbed feature based technique, which is applied by the computing devices, and generates intermediate results using third party analytics agents that are locally configured in the location of the corresponding data repositories, thereby providing the advantages of:
• Maintaining privacy and confidentiality of the user data.
• Maintaining user data in a locked-down location by minimizing data movement across the communication network, which ensures data security, and also conserves bandwidth as only intermediate results are transmitted to the trust index determining server.
The present disclosure generates the prediction result by using a machine learning model which is trained to generate the prediction result based on an optimal feature set, thereby providing the advantages of:
• Enabling real-time digital trust services, by reducing the amount of data to be processed for the prediction, which in turn reduces the processing complexity and enhances computation speed.
• Achieving complete automation of digital identity trust index determination, which results in accurate digital trust comprehension as compared to conventional human interpretation of records.
Overall, the present disclosure determines the digital identity trust index for wide variety of applications, accurately and in real-time, while maintaining user privacy and data security.
A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary a variety of optional components are described to illustrate the wide variety of possible embodiments of the invention. When a single device or article is described herein, it will be apparent that more than one device/article (whether or not they cooperate) may be used in place of a single device/article. Similarly, where more than one device or article is described herein (whether or not they cooperate), it will be apparent that a single device/article may be used in place of the more than one device or article or a different number of devices/articles may be used instead of the shown number of devices or programs. The functionality and/or the features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality/features. Thus, other embodiments of the invention need not include the device itself.
The specification has described a method and a system for determining digital identity trust index in digital ecosystems. The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that on-going technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein. Such alternatives fall within the scope of the disclosed embodiments. Also, the words "comprising," "having," "containing," and "including," and other similar forms are intended to be equivalent in meaning and be open-ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items, or meant to be limited to only the listed item or items. It must also be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise.
Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based here on. Accordingly, the embodiments of the present invention are intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.
Referral numerals
Reference Number Description
100 Architecture
101 Plurality of data repositories
103 Plurality of computing devices
105 Plurality of third party analytics agents
107 Trust index determining server
109 Processor of the computing device
111 I/O interface of the computing device
113 Memory of the computing device
203 Data of the computing device
205 Modules of the computing device
207 Contextual data
209 Secure sensitive contextual data
211 Intermediate result data
213 Other data
221 Receiving module of the computing device
223 Data retrieving module
225 Secure data generating module
227 Intermediate result generating module
229 Transmitting module
231 Other modules of the computing device
235 Data of the trust index determining server
237 Modules of the trust index determining server
239 Predicted data
241 Other data of the trust index determining server
243 Receiving module of the trust index determining server
245 Feature selection module
247 Trust index determining module
249 Trust index updating module
251 Other modules of the trust index determining server
253 Exemplary subscriber
255 Exemplary financial entity
400 Exemplary computer system
401 I/O Interface of the exemplary computer system
402 Processor of the exemplary computer system
403 Network interface
404 Storage interface
405 Memory of the exemplary computer system
406 User interface
407 Operating system
408 Web browser
409 Communication network
411 Input devices
412 Output devices
| # | Name | Date |
|---|---|---|
| 1 | 202041037553-STATEMENT OF UNDERTAKING (FORM 3) [31-08-2020(online)].pdf | 2020-08-31 |
| 2 | 202041037553-REQUEST FOR EXAMINATION (FORM-18) [31-08-2020(online)].pdf | 2020-08-31 |
| 3 | 202041037553-POWER OF AUTHORITY [31-08-2020(online)].pdf | 2020-08-31 |
| 4 | 202041037553-FORM 18 [31-08-2020(online)].pdf | 2020-08-31 |
| 5 | 202041037553-FORM 1 [31-08-2020(online)].pdf | 2020-08-31 |
| 6 | 202041037553-DRAWINGS [31-08-2020(online)].pdf | 2020-08-31 |
| 7 | 202041037553-DECLARATION OF INVENTORSHIP (FORM 5) [31-08-2020(online)].pdf | 2020-08-31 |
| 8 | 202041037553-COMPLETE SPECIFICATION [31-08-2020(online)].pdf | 2020-08-31 |
| 9 | 202041037553-FORM-26 [16-09-2020(online)].pdf | 2020-09-16 |
| 10 | 202041037553-Proof of Right [28-10-2020(online)].pdf | 2020-10-28 |
| 11 | Abstract 202041037553.jpg | 2021-10-18 |
| 12 | 202041037553-FER.pdf | 2022-03-07 |
| 13 | 202041037553-FER_SER_REPLY [04-07-2022(online)].pdf | 2022-07-04 |
| 14 | 202041037553-CORRESPONDENCE [04-07-2022(online)].pdf | 2022-07-04 |
| 15 | 202041037553-CLAIMS [04-07-2022(online)].pdf | 2022-07-04 |
| 16 | 202041037553-ABSTRACT [04-07-2022(online)].pdf | 2022-07-04 |
| 17 | 202041037553-US(14)-HearingNotice-(HearingDate-27-05-2024).pdf | 2024-04-26 |
| 1 | __search1E_07-03-2022.pdf |