e1b17e075b527db8c6f054f8157df7bb7db07561 - asterixdb

commit	e1b17e075b527db8c6f054f8157df7bb7db07561	[log] [tgz]
author	Ian Maxon <ian@maxons.email>	Tue Sep 17 18:37:48 2024 -0700
committer	Ian Maxon <imaxon@apache.org>	Wed Nov 27 03:30:40 2024 +0000
tree	eb5f5233382f22a039501cc247eb23b8a6e7cd78
parent	dec9f981e425ff749d74df9e08120c71c867a7ba [diff]

[ASTERIXDB-3477,3474][RT] Profiling fixes and improvements

- user model changes: yes
- storage format changes: no
- interface changes: no

Details:

The operator id was being improperly used as the key when updating
the min and max time. It needs to include the activity id to clearly
differentiate between the different activities of an operator.

Exchange time was previously being reported as the difference between
the open and close time of the connector, which is not right. Calculating
the actual time a connector is taking will require careful consideration
of each connector type and usage, such as 1-N vs N-M and pipelined vs.
materializing. Until then it is better to simply omit it, because it is
not typically a predominant factor in query time  If it is, the sampled
partition profile still exists and will be useful to find those cases.

Previously, operators that were the last before a non 1-1 exchange would
not have time or cardinalities reported. This is now fixed.

All operators and connectors now report total cardinalities aggregated
across all partitions.

Ext-ref:MB-63566
Change-Id: I172f0044112777aec7200a0c6ae906711fcdc5f2
Reviewed-on: https://asterix-gerrit.ics.uci.edu/c/asterixdb/+/18905
Reviewed-by: Ali Alsuliman <ali.al.solaiman@gmail.com>
Integration-Tests: Jenkins <jenkins@fulliautomatix.ics.uci.edu>
Tested-by: Jenkins <jenkins@fulliautomatix.ics.uci.edu>

15 files changed

tree: eb5f5233382f22a039501cc247eb23b8a6e7cd78

README.md

What is AsterixDB?

AsterixDB is a BDMS (Big Data Management System) with a rich feature set that sets it apart from other Big Data platforms. Its feature set makes it well-suited to modern needs such as web data warehousing and social data storage and analysis. AsterixDB has:

Data model
A semistructured NoSQL style data model (ADM) resulting from extending JSON with object database ideas
Query languages
An expressive and declarative query language (SQL++ that supports a broad range of queries and analysis over semistructured data
Scalability
A parallel runtime query execution engine, Apache Hyracks, that has been scale-tested on up to 1000+ cores and 500+ disks
Native storage
Partitioned LSM-based data storage and indexing to support efficient ingestion and management of semistructured data
External storage
Support for query access to externally stored data (e.g., data in HDFS) as well as to data stored natively by AsterixDB
Data types
A rich set of primitive data types, including spatial and temporal data in addition to integer, floating point, and textual data
Indexing
Secondary indexing options that include B+ trees, R trees, and inverted keyword (exact and fuzzy) index types
Transactions
Basic transactional (concurrency and recovery) capabilities akin to those of a NoSQL store

Learn more about AsterixDB at its website.

Build from source

To build AsterixDB from source, you should have a platform with the following:

A Unix-ish environment (Linux, OS X, will all do).
git
Maven 3.3.9 or newer.
JDK 11 or newer.
Python 3.6+ with pip and venv

Instructions for building the master:

Checkout AsterixDB master:

  $git clone https://github.com/apache/asterixdb.git

Build AsterixDB master:

  $cd asterixdb
  $mvn clean package -DskipTests

Run the build on your machine

Here are steps to get AsterixDB running on your local machine:

Start a single-machine AsterixDB instance:

  $cd asterixdb/asterix-server/target/asterix-server-*-binary-assembly/apache-asterixdb-*-SNAPSHOT
  $./opt/local/bin/start-sample-cluster.sh

Good to go and run queries in your browser at:
```
  http://localhost:19006
```
Read more documentation to learn the data model, query language, and how to create a cluster instance.

Documentation

To generate the documentation, run asterix-doc with the generate.rr profile in maven, e.g mvn -Pgenerate.rr ... Be sure to run mvn package beforehand or run mvn site in asterix-lang-sqlpp to generate some resources that are used in the documentation that are generated directly from the grammar.

master | 0.9.7 | 0.9.6 | 0.9.5 | 0.9.4.1 | 0.9.4 | 0.9.3 | 0.9.2 | 0.9.1 | 0.9.0

Community support

Users maling list: users@asterixdb.apache.org Join the list by sending an email to users-subscribe@asterixdb.apache.org
Developers and contributors mailing list:dev@asterixdb.apache.org Join the list by sending an email to dev-subscribe@asterixdb.apache.org