TL;DR: In this article, a data acquisition and perusal system and method including a database selection module, a database index generator module and a search module is presented, which allows for the capture of HTML data which is automatically indexed without human intervention and has the ability to automatically and accurately locate or "pinpoint" and highlight specific text or groups of text designated by the user within the resulting database.
Abstract: A data acquisition and perusal system and method including a database selection module, a database index generator module and a search module. The database selection module enables selection of a plurality of files for inclusion into at least one selectable database. The database index generator module enables generation of a searchable index of the data contained in the selectable database. The search module enables a search to be performed of the searchable index according to search criteria. The system allows for the capture of HTML data which is automatically indexed without human intervention and has the ability to automatically and accurately locate or “pinpoint,” and highlight specific text or groups of text designated by the user within the resulting database.
TL;DR: In this paper, a virtual database (VDB) is created by creating a set of files in the data storage system, which are linked to the database blocks on the database storage system associated with a point-in-time copy of the source database.
Abstract: Information from multiple databases is retrieved and stored on a database storage system. Multiple point-in-time copies are obtained for each database. A point-in-time copy retrieves data changed in the database since the retrieval of a previous point-in-time copy. A virtual database (VDB) is created by creating a set of files in the data storage system. Each file in the set of files created for a VDB is linked to the database blocks on the database storage system associated with a point-in-time copy of the source database. The set of files associated with the VDB are mounted on a database server allowing the database server to read from and write to the set of files. Workflows based on VDBs allow various usage scenarios based on databases to be implemented efficiently, for example, testing and development, backup and recovery, and data warehouse building.
TL;DR: This work proposes an approach that applies DSE to generate tests for a database application that produces and uses a mock database in test generation, and demonstrates that it can generate tests with higher code coverage than conventional DSE-based techniques.
Abstract: Software testing has been commonly used in assuring the quality of database applications. It is often prohibitively expensive to manually write quality tests for complex database applications. Automated test generation techniques, such as Dynamic Symbolic Execution (DSE), have been proposed to reduce human efforts in testing database applications. However, such techniques have two major limitations: (1) they assume that the database that the application under test interacts with is accessible, which may not always be true; and (2) they usually cannot create necessary database states as a part of the generated tests.To address the preceding limitations, we propose an approach that applies DSE to generate tests for a database application. Instead of using the actual database that the application interacts with, our approach produces and uses a mock database in test generation. A mock database mimics the behavior of an actual database by performing identical database operations on itself. We conducted two empirical evaluations on both a medical device and an open source software system to demonstrate that our approach can generate, without producing false warnings, tests with higher code coverage than conventional DSE-based techniques.
TL;DR: This book addresses issues related to managing data across a distributed database system and gives implementers guidance on hiding discrepancies across systems and creating the illusion of a single repository for users.
Abstract: This book addresses issues related to managing data across a distributed database system. It is unique because it covers traditional database theory and current research, explaining the difficulties in providing a unified user interface and global data dictionary. The book gives implementers guidance on hiding discrepancies across systems and creating the illusion of a single repository for users. It also includes three sample frameworksimplemented using J2SE with JMS, J2EE, and Microsoft .Netthat readers can use to learn how to implement a distributed database management system. IT and development groups and computer sciences/software engineering graduates will find this guide invaluable.
TL;DR: A criterion specifically tailored for SQL queries (SQLFpc), based on Masking Modified Condition Decision Coverage (MCDC) or Full Predicate Coverage and takes into account a wide range of the syntax and semantics of SQL, including selection, joining, grouping, aggregations, subqueries, case expressions and null values is presented.
TL;DR: In this article, a virtual database (VDB) is created by creating a set of files in the data storage system, which are linked to the database blocks on the database storage system associated with a point-in-time copy of the source database.
Abstract: Information from multiple databases is retrieved and stored on a database storage system. Multiple point-in-time copies are obtained for each database. A point-in-time copy retrieves data changed in the database since the retrieval of a previous point-in-time copy. A virtual database (VDB) is created by creating a set of files in the data storage system. Each file in the set of files created for a VDB is linked to the database blocks on the database storage system associated with a point-in-time copy of the source database. The set of files associated with the VDB are mounted on a database server allowing the database server to read from and write to the set of files. Workflows based on VDBs allow various usage scenarios based on databases to be implemented efficiently, for example, testing and development, backup and recovery, and data warehouse building.
TL;DR: In this paper, a cluster manager is configured to manage a plurality of copies of a mid-tier database as a cluster, and the cluster manager may concurrently manage a backend database system.
Abstract: A cluster manager is configured to manage a plurality of copies of a mid-tier database as a mid-tier database cluster. The cluster manager may concurrently manage a backend database system. The cluster manager is configured to monitor for and react to failures of mid-tier database nodes. The cluster manager may react to a mid-tier database failure by, for example, assigning a new active node, creating a new standby node, creating new copies of the mid-tier databases, implementing new replication or backup schemes, reassigning the node's virtual address to another node, or relocating applications that were directly linked to the mid-tier database to another host. Each node or an associated agent may configure the cluster manager to behave in this fashion during initialization, based on common cluster configuration information. Each copy of the mid-tier database may be, for example, a memory resident database. Thus, a node must reload the entire database into memory to recover a copy of the database.
TL;DR: In this article, the authors describe techniques for creating a second snapshot of a first snapshot of data, modifying the first snapshot, and reverting the modifications to the original snapshot, which enable the ability to put the database into multiple states as the database existed at multiple points in time.
Abstract: This application describes techniques for creating a second snapshot of a first snapshot of a set of data, modifying the first snapshot, and reverting the modifications to the first snapshot. For example, portions of one or more transaction logs may be played into a database to put the database in a particular state a particular point in time. The second snapshot may then be used to revert to a prior state of the database such that additional transaction logs may be played into the database. These techniques enable the ability to put the database into multiple states as the database existed at multiple points in time. Therefore, data can be recovered from the database as the data existed at different points in time. Moreover, individual data objects in the database can be accessed and analyzed as the individual data objects existed at different points in time.
TL;DR: In this article, the authors present a method for storing and replicating data of a primary database system in order to support simultaneous queries to the primary and replication database systems, and roll-back the replication data in the at least one replication database system based on the information on transactions processed by the primary database systems and the database lock data.
Abstract: Embodiments of the present disclosure are directed to a database system and methods for storing and replicating data of a primary database. A method for replicating data item may include receiving replication data from a primary database system and replicating the one or more data items of the primary database system in accordance with the replication data. The replication data may include a transaction log including information on transactions processed by the primary database system and database lock data relating to at least one lock on the one or more data items of the primary database system in order to support simultaneous queries to the primary and replication database systems. The method may also include rolling-back the replication data in the at least one replication database system based on the information on transactions processed by the primary database system and the database lock data.
TL;DR: In this article, an application-aware management interface is used to provide a database interface to the data storage system, based on the database interface, database queries are accepted at the database system.
Abstract: A method is used in managing data storage for databases based on application awareness. Content is received via any of a file system interface, a block based interface, and an object based interface to a data storage system, which executes a general purpose operating system. Based on software installed on the data storage system and running on the general purpose operating system, an application-aware management interface is used to provide a database interface to the data storage system. Based on the database interface, database queries are accepted at the data storage system.
TL;DR: This paper introduces MyBenchmark, an offline data generation tool that takes a set of queries as input and generates database instances for which the users can control the characteristics of the resulting workload.
Abstract: To evaluate the performance of database applications and DBMSs, we usually execute workloads of queries on generated databases of different sizes and measure the response time. This paper introduces MyBenchmark, an offline data generation tool that takes a set of queries as input and generates database instances for which the users can control the characteristics of the resulting workload. Applications of MyBenchmark include database testing, database application testing, and application-driven benchmarking. We present the architecture and the implementation algorithms of MyBenchmark. We also present the evaluation results of MyBenchmark using TPC workloads.
TL;DR: An approach for the automatic generation of a test database for a set of SQL queries using a test criterion specifically tailored for the SQL language (SQLFpc) is proposed, generating a testdatabase of reduced size with an elevated coverage and mutation score.
Abstract: Populating test databases with meaningful test data is a difficult task as it involves generating data for many joined tables that must be diverse enough to be able to reveal faults and small enough to make the testing process efficient. This paper proposes an approach for the automatic generation of a test database for a set of SQL queries using a test criterion specifically tailored for the SQL language (SQLFpc). Given as input a schema database and a set of test requirements derived from the application of the test criterion to the target queries, the approach returns a database instance which satisfies the test requirements. Both the schema and the test requirements are modeled in the Alloy language, after which the analyzer generates the test database. The approach is evaluated on a real case study and the results show its feasibility, generating a test database of reduced size with an elevated coverage and mutation score.
TL;DR: In this paper, the authors propose a method of optimizing the interaction between an application or database and a database server, comprising the steps of: (a) routing data between the application and database and the database server through an optimisation system; (b) the optimization system analysing the data and applying rules to the data to speed up the interaction.
Abstract: A method of optimizing the interaction between an application or database and a database server, comprising the steps of: (a) routing data between the application or database and the database server through an optimisation system; (b) the optimisation system analysing the data and applying rules to the data to speed up the interaction between the application or database and the database server.
TL;DR: This paper investigates a new method to utilise user feedback as a way of monitoring changes in the wireless environment, based on a system of “points” given to each AP in the database, which regulates the size of the database.
Abstract: Wi-Fi fingerprinting is a technique which can provide location in GPS-denied environments, relying exclusively on Wi-Fi signals. It first requires the construction of a database of “fingerprints”, i.e. signal strengths from different access points (APs) at different reference points in the desired coverage area. The location of the device is then obtained by measuring the signal strengths at its location, and comparing it with the different reference fingerprints in the database. The main disadvantage of this technique is the labour required to build and maintain the fingerprints database, which has to be rebuilt every time a significant change in the wireless environment occurs, such as installation or removal of new APs, changes in the layout of a building, etc. This paper investigates a new method to utilise user feedback as a way of monitoring changes in the wireless environment. It is based on a system of “points” given to each AP in the database. When an AP is switched off, the number of points associated with that AP will gradually reduce as the users give feedback, until it is eventually deleted from the database. If a new AP is installed, the system will detect it and update the database with new fingerprints. Our proposed system has two main advantages. First it can be used as a tool to monitor the wireless environment in a given place, detecting faulty APs or unauthorised installation of new ones. Second, it regulates the size of the database, unlike other systems where feedback is only used to insert new fingerprints in the database.
TL;DR: In this article, the authors provide a manner of software asset management involving inputting data pertaining to software into a database, performing software product mapping for software product data input to the database, and determining software compliance as a function of data input and mappings that are performed on the data.
Abstract: As aspects of the present invention provides a manner of software asset management involving inputting data pertaining to software into a database; performing software product mapping for software product data input to the database; performing usage rights rule building for usage rights data input to the database; performing software title mapping for software title data input to the database; and determining software compliance as a function of data input to the database and mappings that are performed on the data. Another aspect of the present invention includes method of inputting or importing data into a database m an IT service management system involving: the system importing the data into temporary storage m the database; the system applying validation rules; the system applying the transformation rules; the user reviewing the processed data; the user modifying the processed data; the user requests that the processed data be committed to the database; the system committing the processed data to records in the database.
TL;DR: Preliminary performance comparisons show a very large (an order of magnitude) performance discrepancy with the Amazon cloud version of the SDSS database, suggesting that much work and coordination needs to occur between cloud service providers and their potential database clients before science databases can successfully and effectively be deployed in the cloud.
Abstract: We report on attempts to put an existing scientific (astronomical) database -- the Sloan Digital Sky Survey (SDSS) science archive [1] - in the cloud. Based on our experience, it is either very frustrating or impossible at this time to migrate an existing, complex SQL Server database into current cloud service offerings such as Amazon (EC2) and Microsoft (SQL Azure). Certainly it is impossible to migrate a large database in excess of a TB, but even with (much) smaller databases, the limitations of cloud services make it very difficult to migrate the data to the cloud without making changes to the schema and settings (for example, inability to migrate a spatial indexing library, and several other user-defined functions and stored procedures) that would invalidate performance comparisons between cloud and on-premise versions. So it is not surprising that our preliminary performance comparisons show a very large (an order of magnitude) performance discrepancy with the Amazon cloud version of the SDSS database. We have also not yet investigated the performance tweaks that could be possible within the cloud.Although we managed to successfully migrate (a subset of) the SDSS catalog database to Amazon EC2, we were not able to access the database in a meaningful way from the outside world. Even though this was advertised as a public dataset on the AWS blog, it was not clear how other users or the public would be able to access this data in a meaningful way, if at all.These difficulties suggest that much work and coordination needs to occur between cloud service providers and their potential database clients before science databases can successfully and effectively be deployed in the cloud. This is true not just for large scientific databases but all databases that make extensive use of advanced database management system (DBMS) features for performance and user convenience.
TL;DR: In this paper, a queued database transaction is sent to the clone database engine for pre-processing while the primary database engine is processing a current database transaction and if data needed to process the queued DB transaction is not currently cached in the computer system, the clone DB engine fetches the data and caches it into a buffer accessible by the primary DB engine.
Abstract: A computer system includes an SSD as swap space for a database management system that comprises a primary database engine and at least one clone database engine. A queued database transaction is sent to the clone database engine for pre-processing while the primary database engine is processing a current database transaction and if data needed to process the queued database transaction is not currently cached in the computer system, the clone database engine fetches the data and caches it into a buffer accessible by the primary database engine. When the primary database engine is ready to process the queued database transaction, it will be able to access the needed data without accessing the SSD, thereby avoiding delays resulting from accessing the SSD.
TL;DR: In this paper, the authors propose to achieve scale-out of the distributed database system that assumes a real-time update to be a requirement and which is achieved by dividing the database system into two or more database domains.
Abstract: It is a purpose of this invention to achieve Scale-Out of the distributed database system that assumes a real-time update to be a requirement and which is achieved by dividing the database system into two or more database domains. This is to achieve handling of even larger scale databases while providing even higher performance. Assuming that the large-scale database system has been distributed to two or more of above-mentioned data base domains, in multi transaction processing with real-time update of the database object across two or more of above-mentioned database domain, this invention is achieved by executing the above-mentioned multi transaction processing to the database meta information storage management part in the database meta information management repository device by applying partition topology technology or replication topology technology for exchange and synchronization of meta information such as status information etc. at even higher speeds.
TL;DR: A database management method for use in a computer for managing a database, including: a step of analyzing a query, a first query for searching the database for compressed data; a step for generating a second query for executing a search of time-series data, based on a response result of the second query as discussed by the authors.
Abstract: In a system manages a plurality of pieces of sensor information in a plant, or the like, it can be reducing an amount of data stored in a database and easily a processing for searching a place of an anomaly and an anomaly cause. A database management method for use in a computer for managing a database, the database management method including: a step of analyzing a query; a step of generating a first inquiry for searching the database for compressed data; a step of generating a second inquiry for executing a search of time-series data; a step of extracting given data from the obtained time-series data, based on a response result of the second inquiry; and a step of generating an output result by extracting data to be output to a client computer from the given data.
TL;DR: In this paper, a multi-tiered replicated process database and corresponding method for supporting replication between tiers are disclosed for T2 database tag data structure data replication between tier two databases, which includes a consolidated database that includes process data replicated from a set of T1 database servers.
Abstract: A multi-tiered replicated process database and corresponding method are disclosed for supporting replication between tiers. The multi-tiered replicated process database comprises a tier one (T1) database server computer including a process history database and a replication service. The replication service includes a set of accumulators. Each accumulator is adapted to render a summary T2 database tag data structure from a set of data values retrieved from the process history database for a specified T1 database tag. The replicated database system also includes a tier two (T2) database server computer comprising a consolidated database that includes process data replicated from a set of T1 database servers. At least a portion of the process data replicated from the set of T1 database servers is summary T2 database tag data rendered by the set of accumulators.
TL;DR: In this article, the authors present a system for verifying data in a distributed database using different automated check operations at different times during the database read and update cycles, including first, second, third, and fourth checks.
Abstract: Systems and methods for verifying data in a distributed database using different automated check operations at different times during the database read and update cycles Various functions may be performed including executing a first check during update operations of the database A second check may also be executed during the update operation of the database, and be implemented as an execution thread of an update daemon A third check may be executed at a time interval between update functions of the update daemon A fourth check may be executed during a time that the database is not being updated Integrity of data in the database may be verified by a computer processor based on the first, second, third, and fourth checks
TL;DR: In this paper, a resource balancer is configured to, for each of a multiple database systems hosting database resources assigned to different user entities, generate a system usage score for each database system based on database usage scores of respective database resources hosted by that database system.
Abstract: Various embodiments of a system and method for diminishing workload imbalance across multiple database systems are described. Embodiments may include a resource balancer configured to, for each of a multiple database systems hosting database resources assigned to different user entities, generate a system usage score for that database system based on database usage scores of respective database resources hosted by that database system. Each usage score of a given database resource may indicate a quantity of work performed by the respective database system to process one or more requests directed to that database resource. The resource balancer may also be configured to generate one or more instructions to move a database resource from a first database system having a first system usage score to a second database system having a smaller system usage score in order to diminish an imbalance of workload across the database systems.
TL;DR: In this paper, the authors present an approach to monitor database system resource consumption over various time periods, in conjunction with scheduled data loading, data export, and query operations, by generating a resource consumption map based on the monitoring and adjusting database system workload throttling.
Abstract: Apparatus, systems, and methods may operate to monitor database system resource consumption over various time periods, in conjunction with scheduled data loading, data export, and query operations. Additional activities may include generating a database system resource consumption map based on the monitoring, and adjusting database system workload throttling to accommodate predicted database system resource consumption based on the resource consumption map and current system loading, prior to the current database resource consumption reaching a predefined critical consumption level. The current system loading may be induced by data loading, data export, or query activity. Other apparatus, systems, and methods are disclosed.
TL;DR: In this issue of Communications, Michael Stonebraker discusses the implications of the CAP theorem on database management system applications that span multiple processing sites.
Abstract: http://cacm.acm.org/blogs/blog-cacm The Communications Web site, http://cacm.acm.org, features more than a dozen bloggers in the BLOG@CACM community. In each issue of Communications, we'll publish selected posts or excerpts. twitter Follow us on Twitter at http://twitter.com/blogCACM Michael Stonebraker discusses the implications of the CAP theorem on database management system applications that span multiple processing sites.
TL;DR: The key insight is that the information collected during previous program executions can be used to systematically generate new queries that will maximize the coverage of the application under test, while guaranteeing that the generated test cases focus on the existing data.
Abstract: A database application differs form regular applications in that some of its inputs may be database queries. The program will execute the queries on a database and may use any result values in its subsequent program logic. This means that a user-supplied query may determine the values that the application will use in subsequent branching conditions. At the same time, a new database application is often required to work well on a body of existing data stored in some large database. For systematic testing of database applications, recent techniques replace the existing database with carefully crafted mock databases. Mock databases return values that will trigger as many execution paths in the application as possible and thereby maximize overall code coverage of the database application.In this paper we offer an alternative approach to database application testing. Our goal is to support software engineers in focusing testing on the existing body of data the application is required to work well on. For that, we propose to side-step mock database generation and instead generate queries for the existing database. Our key insight is that we can use the information collected during previous program executions to systematically generate new queries that will maximize the coverage of the application under test, while guaranteeing that the generated test cases focus on the existing data.
TL;DR: This work proposes that users of personal data management applications should be able to create and refine data structures in an ad-hoc way over time, thereby "organically" growing their schemas, and develops a spreadsheet-like direct manipulation interface for this purpose.
Abstract: Non-technical users are increasingly adding structures to their data. This gives rise to the need for database design. However, traditional database design is deliberate and heavy-weight, requiring technical expertise that everyday users may not possess. For this reason, we propose that users of personal data management applications should be able to create and refine data structures in an ad-hoc way over time, thereby "organically" growing their schemas. For this purpose, we develop a spreadsheet-like direct manipulation interface. We show how integrity constraints can still provide value, even in this scenario of frequent schema and data modifications. We also develop a back-end database implementation to support this interface, with a design that permits schema changes at a low cost.We have folded these ideas into a system, called CRIUS, which supports a nested data model and a graphical user interface. From the user's perspective, the chief advantages of CRIUS are its support for simple schema definition and modification through an intuitive drag-and-drop interface, as well as its guidance towards user data entry based on incrementally updated data integrity. We have evaluated CRIUS by means of user studies and performance studies. The user studies indicate that 1) CRIUS makes it much easier for users to design a database, as compared to state-of-the-art GUI database design tools, and 2) CRIUS makes user data entry more efficient and less error-prone. The performance experiments show that 1) the incremental integrity update in CRIUS is very efficient, making the data entry guidance applicable and 2) the backend database implementation in CRIUS significantly improves the performance of schema update tasks, without a significant impact on other operations.
TL;DR: In this article, a database system providing high performance database versioning is described, where a method for restoring databases to a consistent version including creating a cache view of a shared cache and performing undo or redo operations on the cache view only when a log sequence number falls within a certain range.
Abstract: A database system providing high performance database versioning is described. In a database system employing a transaction log, a method for restoring databases to a consistent version including creating a cache view of a shared cache and performing undo or redo operations on the cache view only when a log sequence number falls within a certain range.