SSL CERTIFICATE

WHY YOU NEED AN SSL CERTIFICATE 


                 Introduction Recent numbers from the U.S. Department of Commerce show that online retail is continuing its rapid growth. However, malicious phishing and pharming schemes and fear of inadequate online security cause online retailers to lose out on business as potential customers balk at doing business online, worrying that sensitive data will be abused or compromised. For e-businesses, the key is to build trust: Running a successful online business requires that your customers trust that your business effectively protects their sensitive information from intrusion and tampering. Installing an SSL Certificate from Starfield Technologies on your e-commerce Web site allows you to secure your online business and build customer confidence by securing all online transactions with up to 256-bit encryption.
                An SSL Certificate on your business’ Web site will ensure that sensitive data is kept safe from prying eyes. With a Starfield Technologies SSL Certificate, customers can trust your site. Before issuing a certificate, Starfield Technologies rigorously authenticates the requestor’s domain control and, in the case of High Assurance SSL Certificates, the identity and, if applicable, the business records of the certificate-requesting entity. The authentication process ensures that customers and business partners can rest assured that a Web site protected with a Starfield Technologies certificate can be trusted.
               A Starfield Technologies SSL Certificate provides the security your business needs and the protection your customers deserve. With a Starfield Technologies SSL Certificate, customers will know that your site is secure. Why You Need a Starfield Technologies SSL Certificate In the rapidly expanding world of electronic commerce, security is paramount. Despite booming Internet sales, widespread consumer fear that Internet shopping is not secure still keeps millions of potential shoppers from buying online.
               Only if your customers trust that their credit card numbers and personal information will be kept safe from tampering can you run a successful online business. For online retailers, securing their shopping sites is paramount. If consumers perceive that their credit card information might be compromised online, they are unlikely to do their shopping on the Internet. A Starfield Technologies SSL Certificate provides an easy, cost-effective and secure means to protect customer information and build trust.
               An SSL Certificate enables Secure Sockets Layer (SSL) encryption of your business’ online transactions, allowing you to build an impenetrable fortress around your customers’ credit card information.
 

Starfield Technologies SSL Certificates offer industry-leading security and versatility: 

1. Fully validated
2. Up to 256-bit encryption
3. One-, two- or three-year validity (Turbo SSL Certificates valid up to 10 years)
4. 99% percent browser recognition
5. Stringent authentication
6. Around-the-clock customer support A Starfield Technologies SSL Certificate helps you build an impenetrable fortress around your customers’ credit card information.

What is an SSL Certificate? 

An SSL certificate is a digital certificate that authenticates the identity of a Web site to visiting browsers and encrypts information for the server via Secure Sockets Layer (SSL) technology. A certificate serves as an electronic “passport” that establishes an online entity’s credentials when doing business on the Web. When an Internet user attempts to send confidential information to a Web server, the user’s browser will access the server’s digital certificate and establish a secure connection.

          Information contained in the certificate includes: n The certificate holder’s name (individual or company)* n The certificate’s serial number and expiration date n Copy of the certificate holder’s public key n The digital signature of the certificate-issuing authority To obtain an SSL certificate, one must generate and submit a Certificate Signing Request (CSR) to a trusted Certification Authority, such as Starfield Technologies, which will authenticate the requestor’s identity, existence and domain registration ownership before issuing a certificate.
          Public and Private Keys When you create a CSR, the Web server software with which the request is being generated, creates two unique cryptographic keys: A public key, which is used to encrypt messages to your (i.e., the certificate holder’s) server and is contained in your certificate, and a private key, which is stored on your local computer and “decrypts” the secure messages so they can be read by your server.
          In order to establish an encrypted link between your Web site and your customer’s Web browser your Web server will match your issued SSL certificate to your private key. Because only the Web server has access to its private key, only the server can decrypt SSL-encrypted data. *High Assurance Certificates only. Turbo SSL Certificates only contain the domain name and no information on the individual or company that purchased the certificate. Enabling Safe and Convenient Online Shopping
           A Starfield Technologies SSL Certificate secures safe, easy and convenient Internet shopping. Once an Internet user enters a secure area — by entering credit card information, e-mail address or other personal data, for example — the shopping site’s SSL certificate enables the browser and Web server to build a secure, encrypted connection. The SSL “handshake” process, which establishes the secure session, takes place discreetly behind the scenes, ensuring an uninterrupted shopping experience for the consumer.
           A “padlock” icon in the browser’s status bar and the “https://” prefix in the URL are the only visible indications of a secure session in progress. By contrast, if a user attempts to submit personal information to an unsecured Web site (i.e., a site that is not protected with a valid SSL certificate), the browser’s built-in security mechanism will trigger a warning to the user, reminding him/her that the site is not secure and that sensitive data might be intercepted by third parties. Faced with such a warning, most Internet users likely will look elsewhere to make a purchase.
           Up to 256-Bit Encryption Starfield Technologies SSL certificates support both industry-standard 128-bit (used by all banking infrastructures to safeguard sensitive data) and high-grade 256-bit SSL encryption to secure online transactions. The actual encryption strength on a secure connection using a digital certificate is determined by the level of encryption supported by the user's browser and the server that the Web site resides on. For example, the combination of a Firefox browser and an Apache 2.X Web server enables up to 256-bit AES encryption with Starfield Technologies certificates.
          Encryption strength is measured in key length — number of bits in the key. To decipher an SSL communication, one needs to generate the correct decoding key. Mathematically speaking, 2n possible values exist for an n-bit key. Thus, 40-bit encryption involves 240 possible values. 128- and 256-bit keys involve a staggering 2128 and 2256 possible combinations, respectively, rendering the encrypted data de facto impervious to intrusion. Even with a brute-force attack (the process of systematically trying all possible combinations until the right one is found) cracking a 128- or 256-bit encryption is computationally unfeasible. Stringent Authentication — A Matter of Trust Before Starfield Technologies issues an SSL Certificate, the applicant’s company or personal information undergoes a rigorous authentication procedure that serves to pre-empt online theft and to verify the domain control and, if applicable, the existence and identity of the requesting entity.
         Only through thorough validation of submitted data can the online customer rest assured that online businesses that utilize SSL certificates from Starfield Technologies indeed are to be trusted. SSL Certificates are only issued to entities whose domain control and, depending on certificate type, business credentials and contact information have been verified. Thus, a Starfield Technologies SSL certificate guarantees that the entity that owns the certificate is who it claims to be and has a legal right to use the domain from which it operates.

Starfield Technologies issues three types of SSL Certificates, each of which relies on authentication of a number of elements: High Assurance Certificate — Corporate: Starfield Technologies will authenticate that:
1. The certificate is being issued to an organization that is currently registered with a government authority.
2. The requesting entity controls the domain in the request.
3. The requesting entity is associated with the organization named in the certificate. High Assurance Certificate — Small Business/Sole Proprietor: Starfield Technologies will authenticate that:
4. The individual named in the certificate is the individual who requested the certificate.
5. The requesting individual controls the domain in the request. Starfield Technologies will authenticate that:
6. The requesting entity controls the domain in the request. Phishing and Pharming — How SSL Can Help Phishing and, recently, pharming pose constant threats to Internet users whose sensitive information is under siege by crackers and other cyber crooks.
         
             An SSL certificate from Starfield Technologies can clip the wings of Internet criminals and help prevent Internet users from being victimized by phishing and pharming schemes when attempting to visit your Web site. Phishing schemes – attempts to steal and exploit sensitive personal information – typically try to trick victims into accessing fraudulent sites that pose as legitimate, trusted entities, such as online businesses and banks. Because perpetrators of such attacks will be using and registering domains that resemble those of the spoofed sites, Starfield Technologies, through its stringent fraud-prevention measures, will detect the schemes and deny certificate requests for suspicious domains.
           More sophisticated than phishing, pharming revolves around the concept of hijacking an Internet Service Provider’s (ISP) domain name server (DNS) entries. When a “pharmer” succeeds in such DNS “poisoning” every computer using that ISP for Internet access is directed to the wrong site when the user types in a URL (e.g., www.ebay.com). SSL certificate technology can help prevent pharming attacks, as well. In essence, a “pharmer” simply will not be able to obtain an SSL certificate from Starfield Technologies, as he/she does not control the domain for which the certificate is requested.
           By protecting your Web site with a Starfield Technologies SSL certificate Internet users that attempt to access a site that poses as yours will be instantly alerted that there is a problem with the supposedly secure connection:
1. No lock icon: Because CAs usually won’t issue a certificate to fraudulent phishing or pharming sites, such sites usually do not use SSL encryption. Internet users, therefore, are alerted by the absence of a padlock icon in their browser’s status bar.
2. Name mismatch error: A pharming site could try to use a certificate issued by a CA for a domain owned by the attacker, but the user’s browser will warn the user that the visited URL does not match the certificate presented by the fake Web server.
3. Untrusted CA: A pharming site might attempt to use a certificate issued by an untrusted CA. In this case, the user’s browser will generate the following warning: “the security certificate was issued by a company you have not chosen to trust.” The alert Internet user will instantly abandon his/her activities/ transactions when presented with such warnings. Thus, a Starfield Technologies SSL certificate provides business owners and wary, savvy Internet users with an effective weapon against phishing, pharming and similar cyber swindles.
Establishing a Secure Connection — How SSL Works An SSL-encrypted connection is established via the SSL “handshake” process, which transpires within seconds — transparently to the end user.

In essence, the SSL “handshake” works thus:
1. When accessing an SSL-secured Web site area, the visitor’s browser requests a secure session from the Web server.
2. The server responds by sending the visitor’s browser its server certificate.
3. The browser verifies that the server’s certificate is valid, is being used by the Web site for which it has been issued, and has been issued by a Certificate Authority that the browser trusts.
4. If the certificate is validated, the browser generates a one-time “session” key and encrypts it with the server’s public key.
5. The visitor’s browser sends the encrypted session key to the server so that both server and browser have a copy.
6. The server decrypts the session key using its private key.
7. The SSL “handshake” process is complete, and an SSL connection has been established.
           A padlock icon appears in the browser’s status bar, indicating that a secure session is under way. Conclusion — The Key to Online Security Demand for reliable online security is increasing. Despite booming online sales many consumers continue to believe that shopping online is less safe than doing so at old-fashioned brick-and-mortar stores.
          The key to establishing a successful online business is to build customer trust. Only when potential customers trust that their credit card information and personal data is safe with your business, will they consider making purchases on the Internet.
         
Thanks For read My Article.Any Query Comment Below.

Data Warehousing

Data Warehousing



1. Data Warehousing

          The question most asked now is, How do I build a data warehouse? This is a question that is not so easy to answer. As you will see in this Artical, there are many approaches to building one. However, at the end of all the research, planning, and architecting, you will come to realize that it all starts with a firm foundation. Whether you are building a large centralized data warehouse, one or more smaller distributed data warehouses (sometimes called data marts), or some combination of the two, you will always come to the point where you must decide on how the data is to be structured.
          This is, after all, one of the most key concepts in data warehousing and what differentiates it from the more typical operational database and decision support application building. That is, you structure the data and build applications around it rather than structuring applications and bringing data to them.
           Data warehouse modeling is a process that produces abstract data models for one or more database components of the data warehouse. It is one part of the overall data warehouse development process, which is comprised of other major processes such as data warehouse architecture, design, and construction. We consider the data warehouse modeling process to consist of all tasks related to requirements gathering, analysis, validation, and modeling. Typically for data warehouse development, these tasks are difficult to separate.This may suggest a rather broad gap between modeling and design activities, which in reality certainly is not the case.
           The separation between modeling and design is done for practical reasons: it is our intention to cover the modeling activities and techniques quite extensively. Some trend-setting authors and data warehouse consultants have taken this point to what we consider to be the extreme. That is, they are presenting what they are calling a totally new approach to data modeling. It is called dimensional data modeling, or fact/dimension modeling. Fancy names have been invented to refer to different types of dimensional models, such as star models and snowflake models. Numerous arguments have been presented against traditional entity-relationship (ER) modeling, when used for modeling data in the data warehouse.
            Rather than taking this more extreme position, we believe that every technique has its area of usability. For example, we do support the many criticisms of ER modeling when considered in a specific context of data warehouse data modeling, and there are also criticisms of dimensional modeling. There are many types of data warehouse applications for which ER modeling is not well suited, especially those that address the needs of a well-identified community of data analysts interested primarily in analyzing their business measures in their business context.
            Likewise, there are data warehouse applications that are not well supported at all by star or snowflake models alone. For example, dimensional modeling is not very suitable for making large, corporatewide data models for a data warehouse. With the changing data warehouse landscape and the need for data warehouse modeling, the new modeling approaches and the controversies surrounding traditional modeling and the dimensional modeling approach all merit investigation. And that is another purpose of this post. Because it presents details of data warehouse modeling processes and techniques, the post can also be used as an initiating for those who want to learn data warehouse modeling.


 2. Data Warehousing Architecture and Implementation Choices 

                 In this post we discuss the architecture and implementation choices available for data warehousing. During the discussions we may use the term data mart. Data marts, simply defined, are smaller data warehouses that can function independently or can be interconnected to form a global integrated data warehouse. However, in this post, unless noted otherwise, use of the term data warehouse also implies data mart. Although it is not always the case, choosing an architecture should be done prior to beginning implementation. The architecture can be determined, or modified, after implementation begins. However, a longer delay typically means an increased volume of rework. And, everyone knows that it is more time consuming and difficult to do rework after the fact than to do it right, or very close to right, the first time. The architecture choice selected is a management decision that will be based on such factors as the current infrastructure, business environment, desired management and control structure, commitment to and scope of the implementation effort, capability of the technical environment the organization employs, and resources available.
            The implementation approach selected is also a management decision, and one that can have a dramatic impact on the success of a data warehousing project. The variables affected by that choice are time to completion, return-on-investment, speed of benefit realization, user satisfaction, potential implementation rework, resource requirements needed at any point-in-time, and the data warehouse architecture selected.

 3. Architecture Choices Selection of an architecture

                    Architecture Choices Selection of an architecture will determine, or be determined by, where the data warehouses and/or data marts themselves will reside and where the control resides. For example, the data can reside in a central location that is managed centrally. Or, the data can reside in distributed local and/or remote locations that are either managed centrally or independently. The architecture choices we consider in this book are global, independent, interconnected, or some combination of all three. The implementation choices to be considered are top down, bottom up, or a combination of both. It should be understood that the architecture choices and the implementation choices can also be used in combinations.
      For example, a data warehouse architecture could be physically distributed, managed centrally, and implemented from the bottom up starting with data marts that service a particular workgroup, department, or line of business.

4. Global Warehouse Architecture 

             A global data warehouse is considered one that will support all, or a large part, of the corporation that has the requirement for a more fully integrated data warehouse with a high degree of data access and usage across departments or lines-of-business. That is, it is designed and constructed based on the needs of the enterprise as a whole. It could be considered to be a common repository for decision support data that is available across the entire organization, or a large subset thereof. Top Down Implementation A top down implementation requires more planning and design work to be completed at the beginning of the project. This brings with it the need to involve people from each of the workgroups, departments, or lines of business that will be participating in the data warehouse implementation. Decisions concerning data sources to be used, security, data structure, data quality, data standards, and an overall data model will typically need to be completed before actual implementation begins. The top down implementation can also imply more of a need for an enterprisewide or corporatewide data warehouse with a higher degree of cross workgroup, department, or line of business access to the data.
            This approach is depicted in with this approach, it is more typical to structure a global data warehouse. If data marts are included in the configuration, they are typically built afterward. And, they are more typically populated from the global data warehouse rather than directly from the operational or external data sources. Bottom Up Implementation A bottom up implementation involves the planning and designing of data marts without waiting for a more global infrastructure to be put in place. This does not mean that a more global infrastructure will not be developed; it will be built incrementally as initial data mart implementations expand. This approach is more widely accepted today than the top down approach because immediate results from the data marts can be realized and used as justification for expanding to a more global implementation. depicts the bottom up approach. In contrast to the top down approach, data marts can be built before, or in parallel with, a global data warehouse. And as the figure shows, data marts can be populated either from a global data warehouse or directly from the operational or external data sources.

 5. Architecting the Data A data warehouse

            Architecting the Data A data warehouse is, by definition, a subject-oriented, integrated, time-variant collection of data to enable decision making across a disparate group of users. One of the most basic concepts of data warehousing is to clean, filter, transform, summarize, and aggregate the data, and then put it in a structure for easy access and analysis by those users. But, that structure must first be defined and that is the task of the data warehouse model. In modeling a data warehouse, we begin by architecting the data. By architecting the data, we structure and locate it according to its characteristics. In this chapter, we review the types of data used in data warehousing and provide some basic hints and tips for architecting that data. We then discuss approaches to developing a data warehouse data model along with some of the considerations. Having an enterprise data model (EDM) available would be very helpful, but not required, in developing the data warehouse data model. For example, from the EDM you can derive the general scope and understanding of the business requirements. The EDM would also let you relate the data elements and the physical design to a specific area of interest. Data granularity is one of the most important criteria in architecting the data. On one hand, having data of a high granularity can support any query. However, having a large volume of data that must be manipulated and managed could be an issue as it would impact response times.
          On the other hand, having data of a low granularity would support only specific queries. But, with the reduced volume of data, you would realize significant improvements in performance.

6. Structuring the Data In structuring the data

         Structuring the Data In structuring the data, for data warehousing, we can distinguish three basic types of data that can be used to satisfy the requirements of an organization: · Real-time data · Derived data · Reconciled data In this section, we describe these three types of data according to usage, scope, and currency. You can configure an appropriate data warehouse based on these three data types, with consideration for the requirements of any particular implementation effort. Depending on the nature of the operational systems, the type of business, and the number of users that access the data warehouse, you can combine the three types of data to create the most appropriate architecture for the data warehouse.

7. Data Modeling

              Data Modeling for a Data Warehouse This chapter provides you with a basic understanding of data modeling, specifically for the purpose of implementing a data warehouse. Data warehousing has become generally accepted as the best approach for providing an integrated, consistent source of data for use in data analysis and business decision making. However, data warehousing can present complex issues and require significant time and resources to implement. This is especially true when implementing on a corporatewide basis. To receive benefits faster, the implementation approach of choice has become bottom up with data marts. Implementing in these small increments of small scope provides a larger return-on-investment in a short amount of time. Implementing data marts does not preclude the implementation of a global data warehouse.
            It has been shown that data marts can scale up or be integrated to provide a global data warehouse solution for an organization. Whether you approach data warehousing from a global perspective or begin by implementing data marts, the benefits from data warehousing are significant. The question then becomes, How should the data warehouse databases be designed to best support the needs of the data warehouse users? Answering that question is the task of the data modeler. Data modeling is, by necessity, part of every data processing task, and data warehousing is no exception. As we discuss this topic, unless otherwise specified, the term data warehouse also implies data mart. We consider two basic data modeling techniques in this book: ER modeling and dimensional modeling. In the operational environment, the ER modeling technique has been the technique of choice. With the advent of data warehousing, the requirement has emerged for a technique that supports a data analysis environment. Although ER models can be used to support a data warehouse environment, there is now an increased interest in dimensional modeling for that task. In this chapter, we review why data modeling is important for data warehousing. Then we describe the basic concepts and characteristics of ER modeling and dimensional modeling.

8. The Process of Data Warehousing

            The Process of Data Warehousing This chapter presents a basic methodology for developing a data warehouse. The ideas presented generally apply equally to a data warehouse or a data mart. Therefore, when we use the term data warehouse you can infer data mart. If something applies only to one or the other, that will be explicitly stated. We focus on the process of data modeling for the data warehouse and provide an extended section on the subject but discuss it in the larger context of data warehouse development. The process of developing a data warehouse is similar in many respects to any other development project. Therefore, the process follows a similar path. What follows is a typical, and likely familiar, development cycle with emphasis on how the different components of the cycle affect your data warehouse modeling efforts.
             It is certainly true that there is no one correct or definitive life cycle for developing a data warehouse. We have chosen one simply because it seems to work well for us. Because our focus is really on modeling, the specific life cycle is not an issue here. What is essential is that we identify what you need to know to create an effective model for your data warehouse environment. There are a number of considerations that must be taken into account as we discuss the data warehouse development life cycle. We need not dwell on them, but be aware of how they affect the development effort and understand how they will affect the overall data warehouse design and model. · The life cycle diagram in  seems to infer a single instance of a data warehouse. Clearly, this should be considered a logical view. That is, there could be multiple physical instances of a data warehouse involved in the environment. As an example, consider an implementation where there are multiple data marts. In this case you would iterate through the tasks in the life cycle for each data mart. This approach, however, brings with it an additional consideration, namely, the integration of the data marts. This integration can have an impact on the physical data, with considerations for redundancy, inconsistency, and currency levels. Integration is also especially important because it can require integration of the data models for each of the data marts as well. If dimensional modeling were being used, the integration might take place at the dimension level. Perhaps there could be a more global model that contains the dimensions for the organization. Then when data marts, or multiple instances of a data warehouse, are implemented, the dimensions used could be subsets of those in the global model. This would enable easier integration and consistency in the implementation. · Data marts can be dependent or independent. In the previous consideration we addressed dependent data marts with their need for integration. Independent data marts are basically smaller in scope data warehouses that are stand-alone. In this case the data models can also be independent, but you must understand that this type of implementation can result in data redundancy, inconsistency, and currency levels. The key message of the life cycle diagram is the iterative nature of data warehouse development. This, more than anything else, distinguishes the life cycle of a data warehouse project from other development projects. Whereas all projects have some degree of iteration, data warehouse projects take iteration to the extreme to enable fast delivery of portions of a warehouse. Thus portions of a data warehouse can be delivered while others are still being developed. In most cases, providing the user with some data warehouse function generates immediate benefits. Delivery of a data warehouse is not typically an all-or-nothing proposition. Because the emphasis of this book is on modeling for the data warehouse, we have left out discussion about infrastructure acquisition. Although this would certainly be part of any typical data warehouse effort, it does not directly impact the modeling process.

 9. Requirements Gathering 

         The traditional development cycle focuses on automating the process, making it faster and more efficient. The data warehouse development cycle focuses on facilitating the analysis that will change the process to make it more effective. Efficiency measures how much effort is required to meet a goal. Effectiveness measures how well a goal is being met against a set of expectations. The requirements identified at this point in the development cycle are used to build the data warehouse model. But, the requirements of an organization change over time, and what is true one day is no longer valid the next. How then, do you know when you have successfully identified the user¢s requirements? Although there is no definitive test, we propose that if your requirements address the following questions, you probably have enough information to begin modeling: · Who (people, groups, organizations) is of interest to the user? · What (functions) is the user trying to analyze? · Why does the user need the data? · When (for what point in time) does the data need to be recorded? · Where (geographically, organizationally) do relevant processes occur? · How do we measure the performance or state of the functions being analyzed? There are many methods for deriving business requirements. In general, these methods can be placed in one of two categories: source-driven requirements gathering and user-driven requirements gathering.

 10. Source-Driven Requirements Gathering 

             Source-driven requirements gathering, as the name implies, is a method based on defining the requirements by using the source data in production operational systems. This is done by analyzing an ER model of source data if one is available or the actual physical record layouts and selecting data elements deemed to be of interest.

 11. User-Driven Requirements Gathering 

            User-driven requirements gathering is a method based on defining the requirements by investigating the functions the users perform. This is usually done through a series of meetings and/or interviews with users. The major advantage to this approach is that the focus is on providing what is needed, rather than what is available. In general, this approach has a smaller scope than the source-driven approach. Therefore, it generally produces a useful data warehouse in a shorter timespan. 11. Data Warehouse Modeling Techniques Data warehouse modeling is the process of building a model for the data that is to be stored in the data warehouse. The model produced is an abstract model, and in this sense, it is a representation of reality, or at least a part of reality which the data warehouse is assumed to support. When considered like this, data warehouse modeling seems to resemble traditional database modeling, which most of us are familiar with in the context of database development for operational applications (OLTP database development). This resemblance should be considered with great care, however, because there are a number of significant differences between data warehouse modeling and OLTP database modeling. These differences impact not only the modeling process but also the modeling techniques to be used.

 12. Selecting a Modeling 

               Tool Modeling for data warehousing is significantly different from modeling for operational systems. In data warehousing, quality and content are more important than retrieval response time. Structure and understanding of the data, for access and analysis, by business users is a base criterion in modeling for data warehousing, whereas operational systems are more oriented toward use by software specialists for creation of applications. Data warehousing also is more concerned with data transformation, aggregation, subsetting, controlling, and other process-oriented tasks that are typically not of concern in an operational system. The data warehouse data model also requires information about both the source data that will be used as input and how that data will be transformed and flow to the target data warehouse databases. Thus, the functions required for data modeling tools for data warehousing data modeling have significantly different requirements from those required for traditional data modeling for operational systems. In this chapter we outline some of the functions that are of importance for data modeling tools to support modeling for a data warehouse. The key functions we cover are: diagram notation for both ER models and dimensional models, reverse engineering, forward engineering, source to target mapping of data, data dictionary, and reporting. We conclude with a list of modeling tools.

 13. Populating the Data Warehouse 

Populating is the process of getting the source data from operational and external systems into the data warehouse and data marts. The data is captured from the operational and external systems, transformed into a usable format for the data warehouse, and finally loaded into the data warehouse or the data mart. Populating can affect the data model, and the data model can affect the populating process.

Hadoop Introduction

Hadoop Introduction 


Hadoop is an Apache open source framework written in java that allows distributed processing of large datasets across clusters of computers using simple programming models.
The Hadoop framework application works in an environment that provides distributed storage and computation across clusters of computers.
Hadoop is designed to scale up from single server to thousands of machines, each offering local computation and storage.

Hadoop Architecture Hadoop has two major layers namely:
(a) Processing/Computation layer (MapReduce), and
(b) Storage layer (Hadoop Distributed File System).

MapReduce MapReduce is a parallel programming model for writing distributed applications devised at Google for efficient processing of large amounts of data (multi-terabyte data-sets), on large clusters (thousands of nodes) of commodity hardware in a reliable, fault-tolerant manner.
The MapReduce program runs on Hadoop which is an Apache open-source framework.

Hadoop Distributed File System The Hadoop Distributed File System (HDFS) is based on the Google File System (GFS) and provides a distributed file system that is designed to run on commodity hardware.
It has many similarities with existing distributed file systems. However, the differences from other distributed file systems are significant.
It is highly fault-tolerant and is designed to be deployed on low-cost hardware.
It provides high throughput access to application data and is suitable for applications having large datasets.
Apart from the above-mentioned two core components,
Hadoop framework also includes the following two modules:
1.Hadoop Common: These are Java libraries and utilities required by other Hadoop modules.
2.Hadoop YARN: This is a framework for job scheduling and cluster resource management.

How Does Hadoop Work?
 It is quite expensive to build bigger servers with heavy configurations that handle large scale processing, but as an alternative,
you can tie together many commodity computers with single-CPU,
as a single functional distributed system and practically, the clustered machines can read the dataset in parallel and provide a much higher throughput.
Moreover, it is cheaper than one high-end server.
So this is the first motivational factor behind using Hadoop that it runs across clustered and low-cost machines. Hadoop runs code across a cluster of computers.

This process includes the following core tasks that Hadoop performs:

1. Data is initially divided into directories and files. Files are divided into uniform sized blocks of 128M and 64M (preferably 128M).
2. These files are then distributed across various cluster nodes for further processing.
3. HDFS, being on top of the local file system, supervises the processing.
4. Blocks are replicated for handling hardware failure.
5. Checking that the code was executed successfully.
6. Performing the sort that takes place between the map and reduce stages.
7. Sending the sorted data to a certain computer.
8. Writing the debugging logs for each job. Advantages of Hadoop.

1. Hadoop framework allows the user to quickly write and test distributed systems.
2. It is efficient, and it automatic distributes the data and work across the machines and in turn, utilizes the underlying parallelism of the CPU cores.
3. Hadoop does not rely on hardware to provide fault-tolerance and high availability (FTHA), rather Hadoop library itself has been designed to detect and handle failures at the application layer.
4. Servers can be added or removed from the cluster dynamically and Hadoop continues to operate without interruption.
5. Another big advantage of Hadoop is that apart from being open source, it is compatible on all the platforms since it is Java based.
Installation Hadoop Hadoop is supported by GNU/Linux platform and its flavors.
Therefore, we have to install a Linux operating system for setting up Hadoop environment.
In case you have an OS other than Linux, you can install a Virtualbox software in it and have Linux inside the Virtualbox.
Pre-installation Setup Before installing Hadoop into the Linux environment, we need to set up Linux using ssh (Secure Shell).
Follow the steps given below for setting up the Linux environment.

Hadoop Downloading Download and extract Hadoop
2.4.1 from Apache software foundation using the following commands.
$ su password: # cd /usr/local
# wget http://apache.claz.org/hadoop/common/hadoop-2.4.1/
hadoop-2.4.1.tar.gz
# tar xzf
hadoop-2.4.1.tar.gz
# mv hadoop-2.4.1/* to hadoop/
# exit

Modes of Hadoop Operation
Once you have downloaded Hadoop, you can operate your Hadoop cluster in one of the three supported modes:
1. Local/Standalone Mode: After downloading Hadoop in your system, by default, it is configured in a standalone mode and can be run as a single java process.
2. Pseudo Distributed Mode: It is a distributed simulation on single machine. Each Hadoop daemon such as hdfs, yarn, MapReduce etc., will run as a separate java process. This mode is useful for development. 3. Fully Distributed Mode: This mode is fully distributed with minimum two or more machines as a cluster. We will come across this mode in detail in the coming chapters. Installing Hadoop in Standalone Mode Here we will discuss the installation of Hadoop

2.4.1 in standalone mode.
There are no daemons running and everything runs in a single JVM. Standalone mode is suitable for running MapReduce programs during development, since it is easy to test and debug them.
Setting Up Hadoop. You can set Hadoop environment variables by appending the following commands to ~/.bashrc file.
export HADOOP_HOME=/usr/local/hadoop Before proceeding further, you need to make sure that Hadoop is working fine.
Just issue the following command: $ hadoop version
If everything is fine with your setup,
then you should see the following result:
Hadoop 2.4.1 Subversion https://svn.apache.org/repos/asf/hadoop/common -r 1529768
Compiled by hortonmu on 2013-10-07T06:28Z Compiled with protoc 2.5.0 From source with checksum 79e53ce7994d1628b240f09af91e1af4
It means your Hadoop's standalone mode setup is working fine. By default, Hadoop is configured to run in a non-distributed mode on a single machine.
HDFS OVERVIEW Hadoop File System was developed using distributed file system design. It is run on commodity hardware.
Unlike other distributed systems, HDFS is highly fault-tolerant and designed using low-cost hardware.
HDFS holds very large amount of data and provides easier access. To store such huge data, the files are stored across multiple machines.
These files are stored in redundant fashion to rescue the system from possible data losses in case of failure.

HDFS also makes applications available to parallel processing. Features of HDFS It is suitable for the distributed storage and processing.
1. Hadoop provides a command interface to interact with HDFS.
2. The built-in servers of namenode and datanode help users to easily check the status of cluster.
3. Streaming access to file system data.
4. HDFS provides file permissions and authentication. HDFS follows the master-slave architecture and it has the following elements.
Namenode The namenode is the commodity hardware that contains the GNU/Linux operating system and the namenode software.
It is a software that can be run on commodity hardware. The system having the namenode acts as the master server and it does the following tasks:
1. Manages the file system namespace.
2. Regulates client’s access to files.
3. It also executes file system operations such as renaming, closing, and opening files and directories.
Datanode The datanode is a commodity hardware having the GNU/Linux operating system and datanode software.
For every node (Commodity hardware/System) in a cluster, there will be a datanode. These nodes manage the data storage of their system.
1. Datanodes perform read-write operations on the file systems, as per client request.
2. They also perform operations such as block creation, deletion, and replication according to the instructions of the namenode.
Block Generally the user data is stored in the files of HDFS. The file in a file system will be divided into one or more segments and/or stored in individual data nodes. These file segments are called as blocks.
In other words, the minimum amount of data that HDFS can read or write is called a Block. The default block size is 64MB, but it can be increased as per the need to change in HDFS configuration.
Goals of HDFS Fault detection and recovery: Since HDFS includes a large number of commodity hardware, failure of components is frequent.

Therefore HDFS should have mechanisms for quick and automatic fault detection and recovery.

Huge datasets: HDFS should have hundreds of nodes per cluster to manage the applications having huge datasets.

Hardware at data: A requested task can be done efficiently, when the computation takes place near the data.
 Especially where huge datasets are involved, it reduces the network traffic and increases the throughput.

What is MapReduce?
 MapReduce is a processing technique and a program model for distributed computing based on java.
The MapReduce algorithm contains two important tasks, namely Map and Reduce. Map takes a set of data and converts it into another set of data, where individual elements are broken down into tuples (key/value pairs).
Secondly, reduce task, which takes the output from a map as an input and combines those data tuples into a smaller set of tuples.
As the sequence of the name MapReduce implies, the reduce task is always performed after the map job.
The major advantage of MapReduce is that it is easy to scale data processing over multiple computing nodes.
Under the MapReduce model, the data processing primitives are called mappers and reducers.
Decomposing a data processing application into mappers and reducers is sometimes nontrivial.
But, once we write an application in the MapReduce form, scaing the application to run over hundreds, thousands, or even tens of thousands of machines in a cluster is merely a configuration change.
This simple scalability is what has attracted many programmers to use the MapReduce model.

The Algorithm

1. Generally MapReduce paradigm is based on sending the computer to where the data resides!
2. MapReduce program executes in three stages, namely map stage, shuffle stage, and reduce stage.
A) Map stage: The map or mapper’s job is to process the input data. Generally the input data is in the form of file or directory and is stored in the Hadoop file system (HDFS).
The input file is passed to the mapper function line by line. The mapper processes the data and creates several small chunks of data.
B) Reduce stage: This stage is the combination of the Shuffle stage and the Reduce stage. The Reducer’s job is to process the data that comes from the mapper.
After processing, it produces a new set of output, which will be stored in the HDFS.

3. During a MapReduce job, Hadoop sends the Map and Reduce tasks to the appropriate servers in the cluster.
4. The framework manages all the details of data-passing such as issuing tasks, verifying task completion, and copying data around the cluster between the nodes.
5. Most of the computing takes place on nodes with data on local disks that reduces the network traffic.
6. After completion of the given tasks, the cluster collects and reduces the data to form an appropriate result, and sends it back to the Hadoop server.
Inputs and Outputs (Java Perspective) The MapReduce framework operates on pairs, that is, the framework views the input to the job as a set of pairs and produces a set of pairs as the output of the job, conceivably of different types.
The key and the value classes should be in serialized manner by the framework and hence, need to implement the Writable interface. Additionally, the key classes have to implement the Writable-Comparable interface to facilitate sorting by the framework.

Input and Output types of a MapReduce job: (Input) -> map -> -> reduce -> (Output). Terminology
1. PayLoad - Applications implement the Map and the Reduce functions, and form the core of the job.
2. Mapper - Mapper maps the input key/value pairs to a set of intermediate key/value pair.
3. NamedNode - Node that manages the Hadoop Distributed File System (HDFS).
4. DataNode - Node where data is presented in advance before any processing takes place.
5. MasterNode - Node where JobTracker runs and which accepts job requests from clients.
6. SlaveNode - Node where Map and Reduce program runs.
7. JobTracker - Schedules jobs and tracks the assign jobs to Task tracker.
8. Task Tracker - Tracks the task and reports status to JobTracker.
9. Job - A program is an execution of a Mapper and Reducer across a dataset.
10. Task - An execution of a Mapper or a Reducer on a slice of data.
11. Task Attempt - A particular instance of an attempt to execute a task on a SlaveNode. Hadoop Distributions Hadoop Distributions aim to resolve version incompatibilities

• Distribution Vendor will – Integration Test a set of Hadoop products – Package Hadoop products in various installation formats
• Linux Packages, tarballs, etc. – Distributions may provide additional scripts to execute Hadoop – Some vendors may choose to backport features and bug fixes made by Apache – Typically vendors will employ Hadoop committers so the bugs they find will make it into Apache’s repository.

Distribution Vendors

• Cloudera Distribution for Hadoop (CDH)
• MapR Distribution
• Hortonworks Data Platform (HDP)
• Apache BigTop Distribution
• Greenplum HD Data Computing Appliance Cloudera Distribution for Hadoop (CDH)
• Cloudera has taken the lead on providing Hadoop Distribution – Cloudera is affecting the Hadoop eco-system in the same way RedHat popularized Linux in the enterprise circles
• Most popular distribution – http://www.cloudera.com/hadoop – 100% open-source
• Cloudera employs a large percentage of core Hadoop committers
• CDH is provided in various formats – Linux Packages, Virtual Machine Images, and Tarballs Cloudera Distribution for Hadoop (CDH)
• Integrates majority of popular Hadoop products – HDFS, MapReduce, HBase, Hive, Mahout, Oozie, Pig, Sqoop, Whirr, Zookeeper, Flume
• CDH4 is used in this class Supported Operating Systems • Each Distribution will support its own list of Operating Systems (OS)
• Common OS supported – Red Hat Enterprise – CentOS – Oracle Linux – Ubuntu – SUSE Linux Enterprise Server
• Please see vendors documentation for supported OS and version – Supported Operating Systems for CDH4: https://ccp.cloudera.com/display/CDH4DOC/Before+You+Install+CDH4+on+a+Cl uster#BeforeYouInstallCDH4onaCluster-SupportedOperatingSystemsforCDH4
Thanks For read My Article.Any Query Comment Below.

Artificial Intelligent-IV

Artificial Intelligent-IV Hello ,                So    we have go forward to learn new about Artificial Intelligent S...