TL;DR: In this paper, a fairly general procedure is studied to perturb a multivariate density satisfying a weak form of multivariate symmetry, and to generate a whole set of non-symmetric densities.
Abstract: Summary. A fairly general procedure is studied to perturb a multivariate density satisfying a weak form of multivariate symmetry, and to generate a whole set of non-symmetric densities. The approach is sufficiently general to encompass some recent proposals in the literature, variously related to the skew normal distribution. The special case of skew elliptical densities is examined in detail, establishing connections with existing similar work. The final part of the paper specializes further to a form of multivariate skew t-density. Likelihood inference for this distribution is examined, and it is illustrated with numerical examples.
TL;DR: It is proved that speech samples during voice activity intervals are Laplacian random variables, and all marginal distributions of speech are accurately described by LD in decorrelated domains.
Abstract: It is demonstrated that the distribution of speech samples is well described by Laplacian distribution (LD). The widely known speech distributions, i.e., LD, Gaussian distribution (GD), generalized GD, and gamma distribution, are tested as four hypotheses, and it is proved that speech samples during voice activity intervals are Laplacian random variables. A decorrelation transformation is then applied to speech samples to approximate their multivariate distribution. To do this, speech is decomposed using an adaptive Karhunen-Loeve transform or a discrete cosine transform. Then, the distributions of speech components in decorrelated domains are investigated. Experimental evaluations prove that the statistics of speech signals are like a multivariate LD. All marginal distributions of speech are accurately described by LD in decorrelated domains. While the energies of speech components are time-varying, their distribution shape remains Laplacian.
TL;DR: In this article, a dependence measure that characterises dependence at the bivariate level, for all pairs and all higher orders up to and including the dimension of the variable, is presented and sufficient conditions for subsets of dependence measures to be self-consistent.
Abstract: We present properties of a dependence measure that arises in the study of extreme values in multivariate and spatial problems. For multivariate problems the dependence measure characterises dependence at the bivariate level, for all pairs and all higher orders up to and including the dimension of the variable. Necessary and sufficient conditions are given for subsets of dependence measures to be self‐consistent, that is to guarantee the existence of a distribution with such a subset of values for the dependence measure. For pairwise dependence, these conditions are given in terms of positive semidefinite matrices and non‐differentiable, positive definite functions. We construct new nonparametric estimators for the dependence measure which, unlike all naive nonparametric estimators, impose these self‐consistency properties. As the new estimators provide an improvement on the naive methods, both in terms of the inferential and interpretability properties, their use in exploratory extreme value analyses should aid the identification of appropriate dependence models. The methods are illustrated through an analysis of simulated multivariate data, which shows that a lack of self‐consistency is frequently a problem with the existing estimators, and by a spatial analysis of daily rainfall extremes in south‐west England, which finds a smooth decay in extremal dependence with distance.
TL;DR: In this article, the authors investigated the possibility of using nonparametric methods to estimate the univariate marginal distributions in each of the products, as well as the mixing proportion in a mixture of two distributions, each having independent components.
Abstract: Suppose k-variate data are drawn from a mixture of two distributions, each having independent components. It is desired to estimate the univariate marginal distributions in each of the products, as well as the mixing proportion. This is the setting of two-class, fully parametrized latent models that has been proposed for estimating the distributions of medical test results when disease status is unavailable. The problem is one of inference
in a mixture of distributions without training data, and until now it has been tackled only in a fully parametric setting. We investigate the possibility of using nonparametric methods. Of course, when k=1 the problem is not identifiable from a nonparametric viewpoint. We show that the problem is "almost" identifiable when k=2; there, the set of all possible representations can be expressed, in terms of any one of those representations, as a two-parameter family. Furthermore, it is proved that when $k\geq3$ the problem is nonparametrically identifiable under particularly mild regularity conditions. In this case we introduce root-n consistent nonparametric estimators of the 2k univariate marginal distributions and the mixing proportion. Finite-sample and asymptotic properties of the estimators are described.
TL;DR: This paper proposed a multivariate binomial probit model for analyzing multiple response data and use standard multivariate analysis techniques to conduct exploratory analysis on the latent multivariate normal distribution, addressing identifying restrictions that lead to the covariance matrix specified with unit-diagonal elements.
Abstract: Multiple response questions, also known as a pick any/J format, are frequently encountered in the analysis of survey data. The relationship among the responses is difficult to explore when the number of response options, J, is large. The authors propose a multivariate binomial probit model for analyzing multiple response data and use standard multivariate analysis techniques to conduct exploratory analysis on the latent multivariate normal distribution. A challenge of estimating the probit model is addressing identifying restrictions that lead to the covariance matrix specified with unit-diagonal elements (i.e., a correlation matrix). The authors propose a general approach to handling identifying restrictions and develop specific algorithms for the multivariate binomial probit model. The estimation algorithm is efficient and can easily accommodate many response options that are frequently encountered in the analysis of marketing data. The authors illustrate multivariate analysis of multiple respo...
TL;DR: A number of families of copula functions are given, with attention focusing on those that fall within the Archimedean class, showing members of this class of copulas to be rich in various distributional attributes that are desired when modelling.
Abstract: Summary By a theorem due to Sklar, a multivariate distribution can be represented in terms of its underlying margins by binding them together using a copula function. By exploiting this representation, the 'copula approach' to modelling proceeds by specifying distributions for each margin and a copula function. In this paper, a number of families of copula functions are given, with attention focusing on those that fall within the Archimedean class. Members of this class of copulas are shown to be rich in various distributional attributes that are desired when modelling. The paper then proceeds by applying the copula approach to construct mod els for data that may suffer from selectivity bias. The models examined are the self-selection model, the switching regime model and the double-selection model. It is shown that when models are constructed using copulas from the Archimedean class, the resulting expressions for the log-likelihood and score facilitate maximum likelihood estimation. The literature on selectivity modelling is almost exclusively based on multivariate normal specifications. The copula approach permits selection modelling based on multivariate non-normality. Examples of self-selection models for labour supply and for duration of hospitalization illustrate the application of the copula approach to modelling.
TL;DR: This article provided numerically reliable analytical expressions for the score, Hessian, and information matrix of conditionally heteroscedastic dynamic regression models when the conditional distribution is multivariatet.
Abstract: We provide numerically reliable analytical expressions for the score, Hessian, and information matrix of conditionally heteroscedastic dynamic regression models when the conditional distribution is multivariatet. We also derive one-sided and two-sided Lagrange multiplier tests for multivariate normality versus multivariate t based on the first two moments of the squared norm of the standardized innovations evaluated at the Gaussian pseudo-maximum likelihood estimators of the conditional mean and variance parameters. Finally, we illustrate our techniques through both Monte Carlo simulations and an empirical application to 26 U.K. sectorial stock returns that confirms that their conditional distribution has fat tails.
TL;DR: In this paper, a nonparametric cumulative sum procedure is proposed to detect possible shifts in the mean vector of a multivariate measurement of a statistical process when the multivariate distribution of the measurement is non-Gaussian.
Abstract: Summary. The fairly limited range of tools for multivariate statistical process control generally rests on the assumption that the data vectors follow a multivariate normal distribution-an assumption that is rarely satisfied. We discuss detecting possible shifts in the mean vector of a multivariate measurement of a statistical process when the multivariate distribution of the measurement is non-Gaussian. A nonparametric cumulative sum procedure is suggested which is based both on the order information among the measurement components and on the order information between the measurement components and their in-control means. It is shown that this procedure is effective in detecting a wide range of possible shifts. Several numerical examples are presented to evaluate its performance. This procedure is also applied to a data set from an aluminium smelter.
TL;DR: An efficient approach for the evaluation of the Nakagami-m (1960) multivariate probability density function (PDF) and cumulative distribution function (CDF) with arbitrary correlation is presented and useful closed formulas are derived.
Abstract: An efficient approach for the evaluation of the Nakagami-m (1960) multivariate probability density function (PDF) and cumulative distribution function (CDF) with arbitrary correlation is presented. Approximating the correlation matrix with a Green's matrix, useful closed formulas for the joint Nakagami-m PDF and CDF, are derived. The proposed approach is a significant theoretical tool that can be efficiently used in the performance analysis of wireless communications systems operating over correlative Nakagami-m fading channels.
TL;DR: It is shown that this approach to clustering can be extended to analyse data with mixed categorical and continuous attributes and where some of the data are missing at random in the sense of Little and Rubin (Statistical Analysis with Mixing Data).
TL;DR: The stepwise conditional transformation technique is proposed to transform multiple variables to be univariateGaussian and multivariate Gaussian with no cross correlation.
Abstract: Most geostatistical studies consider multiple-related variables These relationships often show complex features such as nonlinearity, heteroscedasticity, and mineralogical or other constraints These features are not handled by the well-established Gaussian simulation techniques Earth science variables are rarely Gaussian Transformation or anamorphosis techniques make each variable univariate Gaussian, but do not enforce bivariate or higher order Gaussianity The stepwise conditional transformation technique is proposed to transform multiple variables to be univariate Gaussian and multivariate Gaussian with no cross correlation This makes it remarkably easy to simulate multiple variables with arbitrarily complex relationships: (1) transform the multiple variables, (2) perform independent Gaussian simulation on the transformed variables, and (3) back transform to the original variables The back transformation enforces reproduction of the original complex features The methodology and underlying assumptions are explained Several petroleum and mining examples are used to show features of the transformation and implementation details
TL;DR: In this article, the authors explore the usefulness of the multivariate skew-normal distribution in the context of graphical models and show how the factorization of the likelihood function according to a graph can be rearranged in order to obtain a parameter based factorization.
Abstract: This paper explores the usefulness of the multivariate skew-normal distribution in the context of graphical models A slight extension of the family recently discussed by Azzalini & Dalla Valle (1996) and Azzalini & Capitanio (1999) is described, the main motivation being the additional property of closure under conditioning After considerations of the main probabilistic features, the focus of the paper is on the construction of conditional independence graphs for skew-normal variables Necessary and sufficient conditions for conditional independence are stated, and the admissible structures of a graph under restriction on univariate marginal distribution are studied Finally, parameter estimation is considered It is shown how the factorization of the likelihood function according to a graph can be rearranged in order to obtain a parameter based factorization
TL;DR: In this article, a new bounded influence estimator is proposed that combines high asymptotic efficiency for normal data, high breakdown point behaviour with contaminated data and computational simplicity for large data sets.
Abstract: Summary. Many geophysical regression problems require the analysis of large (more than 104 values) data sets, and, because the data may represent mixtures of concurrent natural processes with widely varying statistical properties, contamination of both response and predictor variables is common. Existing bounded influence or high breakdown point estimators frequently lack the ability to eliminate extremely influential data and/or the computational efficiency to handle large data sets. A new bounded influence estimator is proposed that combines high asymptotic efficiency for normal data, high breakdown point behaviour with contaminated data and computational simplicity for large data sets. The algorithm combines a standard M-estimator to downweight data corresponding to extreme regression residuals and removal of overly influential predictor values (leverage points) on the basis of the statistics of the hat matrix diagonal elements. For this, the exact distribution of the hat matrix diagonal elements pii for complex multivariate Gaussian predictor data is shown to be β(pii, m, N−m), where N is the number of data and m is the number of parameters. Real geophysical data from an auroral zone magnetotelluric study which exhibit severe outlier and leverage point contamination are used to illustrate the estimator's performance. The examples also demonstrate the utility of looking at both the residual and the hat matrix distributions through quantile–quantile plots to diagnose robust regression problems.
TL;DR: Space-varying regression models as mentioned in this paper are generalizations of standard linear models where the regression coefficients are allowed to change in space, and the spatial structure is specified by a multivariate extension of pairwise difference priors.
TL;DR: In this paper, a unified treatment and a Bayesian interpretation of two different classes of multivariate skew-normal distributions proposed by Azzalini and Dalla Valle (Biometrika 83 (1996) 715) and Gupta et al. (2001) is presented.
TL;DR: Incomplete data and the generation mechanisms type of incomplete data and its analysis statistical models for incomplete data analysis of data with missing values missing data in multinominal data algorithms for MLE for multivariate normal datawith missing values scoring method EM algorithm.
Abstract: Incomplete data and the generation mechanisms type of incomplete data and its analysis statistical models for incomplete data analysis of data with missing values missing data in multinominal data algorithms for MLE for multivariate normal datawith missing values scoring method EM algorithm basics of EM algorithm extension of EM algorithm and acceleration EM algorithm as an optimization tool robust model and outlier detection scale mixture model of normal distributions multivariate andcontaminated normal distribution robust tobit model robust factor model statistical model with latent variables latent structure model and EM algorithm latent class model and latent trait model structured equations model with latent variables extensions of EM algorithm ECM algorithm ECME algorithm optimal EM algorithm MCEM algorithm covergence speed of EM algorithm convergence speed comparisons of EM and other optimization algorithms quasi Newton method acceleration methods of the EMalgorithm neural networks and EM algorithm EM algorithm in neural networks geometric interpretation of EM algorithm Marcov chain Monte Carlo Bayes estimation Marcov chain Metropolis-Hastings algorithm data augmentation algorithm poor man's dataaugmentation algorithm Gibbs sampling algorithm Appendices: SOLAS for missing data analysis Lem
TL;DR: In this article, the semiparametric estimation of multivariate long-range dependent processes is analyzed, and the proposed estimator is shown to have a smaller limiting variance than the two-step Gaussian estimator studied by Lobato (1999).
Abstract: This paper analyzes the semiparametric estimation of multivariate long-range dependent processes. The class of spectral densities considered includes multivariate
fractionally integrated processes, which are not covered by the existing literature. This paper also establishes the consistency of the multivariate Gaussian semiparametric estimator, which has not been shown in the other works.
Asymptotic normality of the multivariate Gaussian semiparametric estimator is also established, and the proposed estimator is shown to have a smaller limiting
variance than the two-step Gaussian semiparametric estimator studied by Lobato (1999). Gaussianity is not assumed in the asymptotic theory.
TL;DR: In this paper, the authors define a general statistical framework for multiple hypothesis testing and show that the correct null distribution for the test statistics is obtained by projecting the true distribution of the test statistic onto the space of mean zero distributions.
Abstract: We define a general statistical framework for multiple hypothesis testing and show that the correct null distribution for the test statistics is obtained by projecting the true distribution of the test statistics onto the space of mean zero distributions. For common choices of test statistics (based on an asymptotically linear parameter estimator), this distribution is asymptotically multivariate normal with mean zero and the covariance of the vector influence curve for the parameter estimator. This test statistic null distribution can be estimated by applying the non-parametric or parametric bootstrap to correctly centered test statistics. We prove that this bootstrap estimated null distribution provides asymptotic control of most type I error rates. We show that obtaining a test statistic null distribution from a data null distribution, e.g. projecting the data generating distribution onto the space of all distributions satisfying the complete null), only provides the correct test statistic null distribution if the covariance of the vector influence curve is the same under the data null distribution as under the true data distribution. This condition is a weak version of the subset pivotality condition. We show that our multiple testing methodology controlling type I error is equivalent to constructing an error-specific confidence region for the true parameter and checking if it contains the hypothesized value. We also study the two sample problem and show that the permutation distribution produces an asymptotically correct null distribution if (i) the sample sizes are equal or (ii) the populations have the same covariance structure. We include a discussion of the application of multiple testing to gene expression data, where the dimension typically far exceeds the sample size. An analysis of a cancer gene expression data set illustrates the methodology.
TL;DR: This paper will focus on techniques used to segment HSI data into homogenous clusters, and the definition of the multivariate Elliptically Contoured Distribution mixture model will be developed.
Abstract: Developing proper models for hyperspectral imaging (HSI) data allows for useful and reliable algorithms for data exploitation. These models provide the foundation for development and evaluation of detection, classification, clustering, and estimation algorithms. To date, real world HSI data has been modeled as a single multivariate Gaussian, however it is well known that real data often exhibits non-Gaussian behavior with multi-modal distributions. Instead of the single multivariate Gaussian distribution, HSI data can be model as a finite mixture model, where each of the mixture components need not be Gaussian. This paper will focus on techniques used to segment HSI data into homogenous clusters. Once the data has been segmented, each individual cluster can be modeled, and the benefits provided by the homogeneous clustering of the data versus non-clustering explored. One of the promising techniques uses the Expectation-Maximization (EM) algorithm to cluster the data into Elliptically Contoured Distributions (ECDs). A larger family of distributions, the family of ECDs includes the mutlivariate Gaussian distribution and exhibits most of its properties. ECDs are uniquely defined by their multivariate mean, covariance and the distribution of its Mahalanobis (or quadratic) distance metric. This metric lets multivariate data be identified using a univariate statistic and can be adjusted to more closely match the longer tailed distributions of real data. This paper will focus on three issues. First, the definition of the multivariate Elliptically Contoured Distribution mixture model will be developed. Second, various techniques will be described that segment the mixed data into homogeneous clusters. Most of this work will focus on the EM algorithm and the multivariate t-distribution, which is a member of the family of ECDs and provides longer tailed distributions than the Gaussian. Lastly, results using HSI data from the AVIRIS sensor will be shown, and the benefits of clustered data will be presented.
TL;DR: In this article, the problem of testing the error distribution in a multivariate linear regression (MLR) model is studied, where empirical multivariate skewness and kurtosis criteria are compared with a simulation-based estimate of their expected value under the hypothesized distribution.
Abstract: We study the problem of testing the error distribution in a multivariate linear regression (MLR) model. The tests are functions of appropriately standardized multivariate least squares residuals whose distribution is invariant to the unknown cross-equation error covariance matrix. Empirical multivariate skewness and kurtosis criteria are then compared with a simulation-based estimate of their expected value under the hypothesized distribution. Special cases considered include testing multivariate normal and stable error distributions. In the Gaussian case, finite-sample versions of the standard multivariate skewness and kurtosis tests are derived. To do this, we exploit simple, double and multi-stage Monte Carlo test methods. For non-Gaussian distribution families involving nuisance parameters, confidence sets are derived for the nuisance parameters and the error distribution. The tests are applied to an asset pricing model with observable risk-free rates, using monthly returns on New York Stock Exchange (NYSE) portfolios over 5-year subperiods from 1926 to 1995.
TL;DR: This paper presents a method using an F0-dependent multivariate normal distribution of which mean is represented by a function of fundamental frequency (FO), which represents the pitch dependency of each feature, while the F1-normalized covariance represents the non-pitch dependency.
Abstract: The pitch dependency of timbres has not been fully exploited in musical instrument identification. In this paper, we present a method using an F0-dependent multivariate normal distribution of which mean is represented by a function of fundamental frequency (FO). This F0-dependent mean function represents the pitch dependency of each feature, while the F0-normalized covariance represents the non-pitch dependency. Musical instrument sounds are first analyzed by the F0-dependent multivariate normal distribution, and then identified by using the discriminant function based on the Bayes decision rule. Experimental results of identifying 6,247 solo tones of 19 musical instruments by 10-fold cross validation showed that the proposed method improved the recognition rate at individual-instrument level from 75.73% to 79.73%, and the recognition rate at category level from 88.20% to 90.65%.
TL;DR: Empirical equations are developed based on the complex multivariate normal distribution that generalize the distributions currently used in maximum-likelihood model and heavy-atom refinement and perform satisfactorily compared with currently used programs.
Abstract: Probabilistic methods involving maximum-likelihood parameter estimation have become a powerful tool in computational crystallography. At the centre of these methods are the relevant probability distributions. Here, equations are developed based on the complex multivariate normal distribution that generalize the distributions currently used in maximum-likelihood model and heavy-atom refinement. In this treatment, the effects of various sources of error in the experiment are considered separately and allowance is made for correlations among sources of error. The multivariate distributions presented are closely related to the distributions previously derived in ab initio phasing and can be applied to many different aspects of a crystallographic structure-determination process including model refinement, density modification, heavy-atom phasing and refinement or combinations of them. The underlying probability distributions for multiple isomorphous replacement are re-examined using these techiques. The re-analysis requires the underlying assumptions to be made explicitly and results in a variance term that, unlike those previously used for maximum-likelihood multiple isomorphous replacement phasing, is expressed explicitly in terms of structure-factor covariances. Test cases presented show that the newly derived multiple isomorphous replacement likelihood functions perform satisfactorily compared with currently used programs.
TL;DR: In this paper, the authors characterize variance vulnerability in terms of two-parameter utility functions and identify the multivariate normal as the only distribution such that the EU-and twoparameter approach are compatible when independent background risks prevail.
Abstract: An agent with two-parameter, mean-variance preferences is called variance vulnerable if an increase in the variance of an exogenous, independent background risk induces the agent to choose a lower level of risky activities. Variance vulnerability resembles the notion of risk vulnerability in the expected utility (EU) framework. First, we characterize variance vulnerability in terms of two-parameter utility functions. Second, we identify the multivariate normal as the only distribution such that EU- and two-parameter approach are compatible when independent background risks prevail. Third, presupposing normality, we show that—analogously to risk vulnerability—temperance is a necessary, and standardness and convex risk aversion are sufficient conditions for variance vulnerability.
TL;DR: In this article, a vector autoregressive (VAR) spectral estimation procedure for constructing heteroskedasticity and autocorrelation consistent (HAC) covariance matrices is proposed.
Abstract: This paper proposes a vector autoregressive (VAR) spectral estimation procedure for constructing heteroskedasticity and autocorrelation consistent (HAC) covariance matrices. We establish the consistency of the VARHAC estimator under general conditions similar to those considered in previous research, and we demonstrate that this estimator converges at a faster rate than the kernel-based estimators proposed by Andrews and Monahan (1992) and Newey and West (1994). In finite samples, Monte Carlo simulation experiments indicate that the VARHAC estimator matches, and in some cases greatly exceeds, the performance of the prewhitened kernel estimator proposed by Andrews and Monahan (1992). These simulation experiments also illustrate several important limitations of kernel-based HAC estimation procedures, and highlight the advantages of explicitly modeling the temporal properties of the error terms.
TL;DR: In this paper, the affine equivariant sign covariance matrix (SCM) introduced by Visuri et al. is shown to be proportional to the inverse of the regular covariance matrices, and an estimate of the covariance and correlation matrix based on the SCM is presented.
TL;DR: In this article, the moments of the autocovariance function and of the sample variogram estimator depend on a measure of multivariate kurtosis, but not on a skewness parameter.
TL;DR: This work derives a simple and theoretically valid approach by establishing the shapes of the sequentially constructed conditional distributions, which ensure histogram reproduction.
TL;DR: It is found that small MIMO systems such as 2 /spl times/ 2 can be considered normally distributed and can also be approximated with a Kronecker structure, while larger systems show evidence of strong non-normality and are not well modeled using a Kr onecker product.
Abstract: Measurements taken at the campus of Brigham Young University (BYU) are used to investigate the statistical properties of the indoor MIMO channel. Two statistical tests, Royston's and Henze-Zirkler's, are applied to the MIMO data to assess whether the data belongs to a multivariate normal distribution or not. The possibility of modeling the covariance matrix as a Kronecker product of the correlations at the transmitter and receiver are also investigated by deriving a likelihood ratio test. It is found that small MIMO systems such as 2 /spl times/ 2 can be considered normally distributed and can also be approximated with a Kronecker structure. Larger systems, on the other hand, show evidence of strong non-normality and are not well modeled using a Kronecker product. However, for short measurement segments, these distributions can be used for approximate channel capacity calculations.
TL;DR: In this paper, weather forecasts can be used in the pricing of weather derivatives and derive results for the most important types of weather index and contract, such as the expected payoff of linear contracts on linear indices.
TL;DR: In this paper, the conditional test procedures for testing elliptical symmetry of multivariate distribution were suggested. But the equivalence between the conditional tests and their unconditional counterparts was not established, and the power behavior of the tests under global as well as local alternatives was investigated theoretically.