TL;DR: GraphPrompt as discussed by the authors unifies pre-training and downstream tasks into a common task template, and employs a learnable prompt to assist a downstream task in locating the most relevant knowledge from the pre-trained model in a task-specific manner.
Abstract: Graphs can model complex relationships between objects, enabling a myriad of Web applications such as online page/article classification and social recommendation. While graph neural networks (GNNs) have emerged as a powerful tool for graph representation learning, in an end-to-end supervised setting, their performance heavily relies on a large amount of task-specific supervision. To reduce labeling requirement, the “pre-train, fine-tune” and “pre-train, prompt” paradigms have become increasingly common. In particular, prompting is a popular alternative to fine-tuning in natural language processing, which is designed to narrow the gap between pre-training and downstream objectives in a task-specific manner. However, existing study of prompting on graphs is still limited, lacking a universal treatment to appeal to different downstream tasks. In this paper, we propose GraphPrompt, a novel pre-training and prompting framework on graphs. GraphPrompt not only unifies pre-training and downstream tasks into a common task template, but also employs a learnable prompt to assist a downstream task in locating the most relevant knowledge from the pre-trained model in a task-specific manner. Finally, we conduct extensive experiments on five public datasets to evaluate and analyze GraphPrompt.
TL;DR: MicrobiotaProcess as discussed by the authors provides a comprehensive data structure, MPSE, to better integrate the primary and intermediate data, which improves the integration and exploration of the downstream data and provides flexible downstream analysis components, and provides visualization methods to assist in presenting and interpreting results.
TL;DR: In this paper , the authors investigated the potential of ChatGPT to aid in clinical text mining by examining its ability to extract structured information from unstructured healthcare texts, with a focus on biological named entity recognition and relation extraction.
Abstract: Recent advancements in large language models (LLMs) have led to the development of highly potent models like OpenAI's ChatGPT. These models have exhibited exceptional performance in a variety of tasks, such as question answering, essay composition, and code generation. However, their effectiveness in the healthcare sector remains uncertain. In this study, we seek to investigate the potential of ChatGPT to aid in clinical text mining by examining its ability to extract structured information from unstructured healthcare texts, with a focus on biological named entity recognition and relation extraction. However, our preliminary results indicate that employing ChatGPT directly for these tasks resulted in poor performance and raised privacy concerns associated with uploading patients' information to the ChatGPT API. To overcome these limitations, we propose a new training paradigm that involves generating a vast quantity of high-quality synthetic data with labels utilizing ChatGPT and fine-tuning a local model for the downstream task. Our method has resulted in significant improvements in the performance of downstream tasks, improving the F1-score from 23.37% to 63.99% for the named entity recognition task and from 75.86% to 83.59% for the relation extraction task. Furthermore, generating data using ChatGPT can significantly reduce the time and effort required for data collection and labeling, as well as mitigate data privacy concerns. In summary, the proposed framework presents a promising solution to enhance the applicability of LLM models to clinical text mining.
Weihua Chen, Xiangmin Xu, Jian Jia, Hao Luo, Yaohua Wang, Fan Wang, Rong Jin, Xiuyu Sun
1 Jun 2023
TL;DR: SOLIDER is a self-supervised learning framework that learns a general human representation from massive unlabeled human images. It utilizes prior knowledge from human images to build pseudo semantic labels and import more semantic information into the learned representation. SOLIDER introduces a conditional network with a semantic controller to produce representations with different ratios of semantic information, which can fit different needs of downstream tasks.
Abstract: Human-centric visual tasks have attracted increasing research attention due to their widespread applications. In this paper, we aim to learn a general human representation from massive unlabeled human images which can benefit downstream human-centric tasks to the maximum extent. We call this method SOLIDER, a Semantic cOntrollable seLf-supervIseD lEaRning framework. Unlike the existing self-supervised learning methods, prior knowledge from human images is utilized in SOLIDER to build pseudo semantic labels and import more semantic information into the learned representation. Meanwhile, we note that different downstream tasks always require different ratios of semantic information and appearance information. For example, human parsing requires more semantic information, while person re-identification needs more appearance information for identification purpose. So a single learned representation cannot fit for all requirements. To solve this problem, SOLIDER introduces a conditional network with a semantic controller. After the model is trained, users can send values to the controller to produce representations with different ratios of semantic information, which can fit different needs of downstream tasks. Finally, SOLIDER is verified on six downstream human-centric visual tasks. It outperforms state of the arts and builds new baselines for these tasks. The code is released in https://github.com/tinyvision/SOLIDER.
TL;DR: In this article , the authors argue that current diversity and inclusion efforts do not prevent exclusiveness, unless the framing and scope of the projects are revisited in public health terms, and they conclude that enhanced focus on socio-environmental determinants of health and aligned public health interventions as precision medicine outputs would be to the benefit of all and especially of those who are most at risk of exclusion.
Abstract: This paper problematizes the precision medicine approach embraced by the All of Us Research Program (US) and by Genomics England (UK) in terms of benefits distribution, by arguing that current "diversity and inclusion" efforts do not prevent exclusiveness, unless the framing and scope of the projects are revisited in public health terms. Grounded on document analysis and fieldwork interviews, this paper analyzes efforts to address potential patterns of exclusion upstream (from participating in precision medicine research) and downstream (from benefitting from precision medicine outputs). It argues that efforts for inclusion upstream are not corresponded downstream, and this unbalance jeopardizes the equitable capacities of the projects. It concludes that enhanced focus on socio-environmental determinants of health and aligned public health interventions as precision medicine outputs would be to the benefit of all and especially of those who are most at risk of (upstream as well as downstream) exclusion.
TL;DR: In this paper , the authors provide an overview of current methods and technological improvements and the latest trends in OGTS to show how this industry strives to achieve sustainable development goals, including increasing flexibility, energy saving, emission reduction, and changing energy structure.
Abstract: The oil & gas transport and storage (OGTS) engineering, from the upstream of gathering and processing in the oil & gas fields, to the midstream long-distance pipelines, and the downstream tanks and LNG terminals, while using supply chains to connect each part, is exploring its way to reduce energy consumption and carbon footprints. This work provides an overview of current methods and technological improvements and the latest trends in OGTS to show how this industry strives to achieve sustainable development goals. The critical analyses are from increasing flexibility, energy saving, emission reduction, and changing energy structure. The study shows the need to focus on improving energy efficiency further, reducing energy/water/material consumption and emissions, and maintaining safety for such an extensive oil & gas network.
TL;DR: In this paper , the authors review the current landscape of commercial health data vendors, with special emphasis on the sources of their data, challenges associated with data reproducibility and generalisability, and ethical considerations for data vending.
TL;DR: The authors measured acoustic, phonetic, and word-level properties encoded in individual layers, using a lightweight analysis tool based on canonical correlation analysis (CCA), and found that these properties evolve across layers differently depending on the model, and the variations relate to the choice of pre-training objective.
Abstract: Many self-supervised speech models, varying in their pre-training objective, input modality, and pre-training data, have been proposed in the last few years. Despite impressive successes on downstream tasks, we still have a limited understanding of the properties encoded by the models and the differences across models. In this work, we examine the intermediate representations for a variety of recent models. Specifically, we measure acoustic, phonetic, and word-level properties encoded in individual layers, using a lightweight analysis tool based on canonical correlation analysis (CCA). We find that these properties evolve across layers differently depending on the model, and the variations relate to the choice of pre-training objective. We further investigate the utility of our analyses for downstream tasks by comparing the property trends with performance on speech recognition and spoken language understanding tasks. We discover that CCA trends provide reliable guidance to choose layers of interest for downstream tasks and that single-layer performance often matches or improves upon using all layers, suggesting implications for more efficient use of pre-trained models. 1
TL;DR: In this paper , a machine-learning approach was introduced to analyze the heterogeneous effect of customer concentration on the bargaining power of downstream customers in supplier-customer interactions, leading to a significant deterioration of financing constraints for upstream firms.
TL;DR: In this paper , the authors examined the effects of PDC on SMEs' operational performance under conditions of environmental turbulence and found that PDC does not directly affect an SME's operational performance.
Abstract:
Purpose
Small- and medium-sized enterprises (SMEs) often operate in environments marked by high levels of turbulence. Such firms adopt digital technologies and platforms that provide access to external real-time information and establish digital connectivity between firms to remain competitive. This study aims to focus on SMEs’ downstream and upstream platform-based digital connectivity (PDC).
Design/methodology/approach
This study examines the effects of PDC on SMEs’ operational performance under conditions of environmental turbulence. The data was gathered from 192 SMEs operating in the manufacturing arena.
Findings
The results show that the adoption of PDC does not directly affect an SME’s operational performance. However, in highly turbulent environments, PDC can improve operational performance. The results indicate that the performance effects of PDC vary according to the level and type of environmental turbulence.
Research limitations/implications
This research offers insights into the relationship between PDC among SMEs and operational performance and encourages future research examining other possible conditional effects that could explain the contradictory results found in previous research.
Originality/value
This study contributes to the knowledge of supply-chain digitalization among SMEs and its performance effects in varying environmental conditions. Further, this study contributes to the prior research by focusing on the interorganizational aspects of digitalization in SMEs.
TL;DR: In this paper , the primary safety factors that contribute to accidents in downstream oil and gas construction projects in Malaysia are identified and evaluated using exploratory factor analysis (EFA) and structural equation modeling (SEM).
TL;DR: In this paper , the impact of demand-pull and technology-push on growth, innovation, and the factor bias of technological change in a two-layer network of input-output (market) and patent citation (innovation) links among 307 6-digit US manufacturing industries in 1977-2012.
Abstract: This paper studies the impact of Demand-pull (DP) and Technology-push (TP) on growth, innovation, and the factor bias of technological change in a two-layer network of input–output (market) and patent citation (innovation) links among 307 6-digit US manufacturing industries in 1977–2012. Two types of TP and DP are distinguished: (1) DP and TP are between-layer spillovers when market demand shocks pull innovation and innovation pushes market growth. (2) Within-layer DP arises if downstream users trigger upstream innovation and growth, while TP effects spill over from up- to downstream industries. The results support between- and within-layer TP: Innovation spillovers from upstream industries drive market growth and innovation. Within the market, upstream supply shocks stimulate growth, but this effect differs across industries. DP is not supported but shows a factor bias favoring labor, while TP comes with a shift towards non-production work. The results are strongest after the 2000s and shed light on the drivers of recent technological change and its factor bias.
TL;DR: In this paper , a downwind-vibrating piezoelectric energy harvester under the disturbance of a downstream baffle (DVPEH) is proposed to offer a promising solution for low structure reliability of traditional PWVEHs under fairly high wind speed.
TL;DR: In this article , the authors focus on the platform entry strategy in a supply chain, where a manufacturer that could voluntarily disclose quality information sells its product via a retailer and find that downstream entry induces the manufacturer to reveal more quality information to the consumer than in the monopoly setting.
TL;DR: In this paper , the authors investigate how a manufacturing plant's internal operations along with its network of connections (upstream and downstream) can have an impact on its recovery time from a disruption.
Abstract: Purpose The purpose of our study is to investigate how a manufacturing plant’s internal operations along with its network of connections (upstream and downstream) can have an impact on its recovery time from a disruption. The authors also examine the inverse-U impact of complexity. Finally, the authors test the moderating role that business continuity management plans (BCP) at the plant level have on recovery time. Design/methodology/approach To test our hypotheses, the authors partnered with Resilinc Corporation, a Silicon Valley-based provider of supply chain risk management solutions to identify focal firms’ suppliers, customers and plant-level data including information on parts, manufacturing activities, bill of materials, alternate sites and formal business continuity plans. The authors employed censored data regression technique (Tobit). Findings Several important findings reveal that the plant’s internal operations and network connections impact recovery time. Specifically, the number of parts manufactured at the plant as well as the number of internal plant processes significantly increase disruption recovery time. In addition, the number of supply chains (upstream and downstream) involving the plant as well as the echelon distance of the plant from its original equipment manufacturer significantly increase recovery time. The authors also find that there exists an inverted-U relationship between complexity and recovery time. Finally, the authors find partial support that BCP will have a negative moderating effect between complexity and recovery time. Originality/value This research highlights gaps in the literature related to supply chain disruption and recovery. There is a need for more accurate methods to measure recovery time, more research on recovery at the supply chain site level and further analysis of the impact of supply chain complexity on recovery time.
TL;DR: In this paper , the dual-credit policy and government subsidy are two fundamental policies affecting the electric vehicle industry in China, and the authors focus on exploring how these two policies have different impacts on R&D intensity of upstream raw material companies, midstream power battery manufacturers, and downstream vehicle manufacturers.
TL;DR: In this paper , the corridor usage of downstream moving fish (6,646 individuals from 42 species) was investigated at four small-scale hydropower plants with different concepts to prevent turbine entrainment and to bypass fish.
Abstract: Introduction: Hydropower plants are frequently equipped with physical and behavioral fish protection barriers to prevent downstream moving fish from harmful turbine passage and to guide them to alternative bypasses. As not only diadromous but also potamodromous fish species migrate and inevitably have to pass hydropower plants, knowledge on corridor usage for a wide range of species is important to identify potential deficits and to improve bypass efficiency. Methods: In this study, the corridor usage of downstream moving fish (6,646 individuals from 42 species) was investigated at four small-scale hydropower plants with different concepts to prevent turbine entrainment and to bypass fish. Results: Despite existing bypasses and fine screens with 15 mm and 20 mm bar spacing to prevent turbine entrainment, a large proportion of fish (35%–88%) still passed the turbines. The mainly poor efficiency of the investigated bypasses was probably due to low discharge and unfavorable bypass location or detectability. The various bypass types were used by a different range of fish species and sizes due to species-specific behavior and differing fish communities between sites. The effectiveness of the investigated downstream corridors was positively correlated with the share of discharge. Discussion: To reduce the negative ecological impacts of hydropower plants on downstream moving fish, well-performing bypasses are required that consider not only current requirements regarding design, dimensioning and location, but also the site-specific fish community. Thus, bypasses should function for the widest possible range of species, which can be achieved through less selective bypass types such as full-depth bypasses, or a combination of different bypass systems. Moreover, less harmful turbine technologies and more effective fish protection systems need to be implemented, since fine screens with 15 mm and 20 mm bar spacing cannot prevent small-bodied fish species and juvenile fish <20 cm from turbine entrainment.
TL;DR: In this article , the authors propose a heterogeneous region embedding with prompt learning (HREP) model, which addresses both intra-region and inter-region correlations through two key modules: Heterogeneous Region Embedding (HRE) and prompt learning for different downstream tasks.
Abstract: The prevalence of region-based urban data has opened new possibilities for exploring correlations among regions to improve urban planning and smart-city solutions. Region embedding, which plays a critical role in this endeavor, faces significant challenges related to the varying nature of city data and the effectiveness of downstream applications. In this paper, we propose a novel framework, HREP (Heterogeneous Region Embedding with Prompt learning), which addresses both intra-region and inter-region correlations through two key modules: Heterogeneous Region Embedding (HRE) and prompt learning for different downstream tasks. The HRE module constructs a heterogeneous region graph based on three categories of data, capturing inter-region contexts such as human mobility and geographic neighbors, and intraregion contexts such as POI (Point-of-Interest) information. We use relation-aware graph embedding to learn region and relation embeddings of edge types, and introduce selfattention to capture global correlations among regions. Additionally, we develop an attention-based fusion module to integrate shared information among different types of correlations. To enhance the effectiveness of region embedding in downstream tasks, we incorporate prompt learning, specifically prefix-tuning, which guides the learning of downstream tasks and results in better prediction performance. Our experiment results on real-world datasets demonstrate that our proposed model outperforms state-of-the-art methods.
TL;DR: A public policy agenda that aims to address inequities related to the well-being of children, creation and perpetuation of residential segregation, and racial segregation can address upstream factors as mentioned in this paper .
Abstract: Policy Points Upstream factors-social structures/systems, cultural factors, and public policy-are primary forces that drive downstream patterns and inequities in health that are observed across race and locations. A public policy agenda that aims to address inequities related to the well-being of children, creation and perpetuation of residential segregation, and racial segregation can address upstream factors. Past successes and failures provide a blueprint for addressing upstream health issues and inhibit health equity.
TL;DR: In this article , the authors discuss a few key instances of molecular antagonistic and interdependent relationships between bacterial secretion systems and their produced functional products and discuss how the intersecretion system functions in bacterial rivalry, virulence, and survival, among other things.
Abstract: Unprecedented insights into the biology and functions of bacteria have been and continue to be gained through studying bacterial secretion systems in isolation. This method, however, results in our understanding of the systems being primarily based on the idea that they operate independently, ignoring the subtleties of downstream interconnections. Gram-negative bacteria are naturally able to adapt to and navigate their frequently varied and dynamic surroundings, mostly because of the covert connections between secretion systems. Therefore, to comprehend some of the linked downstream repercussions for organisms that follow this discourse, it is vital to have mechanistic insights into how the intersecretion system functions in bacterial rivalry, virulence, and survival, among other things. To that purpose, this paper discusses a few key instances of molecular antagonistic and interdependent relationships between bacterial secretion systems and their produced functional products.
TL;DR: In this paper , a greener approach for extraction of PHAs in comparison to methods using hazardous solvent is proposed, which is shown to be more efficient in terms of energy consumption and energy efficiency.
TL;DR: Huang et al. as mentioned in this paper proposed an offsite-tuning framework that can adapt billion-parameter foundation models to downstream data without access to the full model by sending a light-weight adapter and a lossy compressed emulator to the data owner, who then fine-tun the adapter on the downstream data with the emulator's assistance.
Abstract: Transfer learning is important for foundation models to adapt to downstream tasks. However, many foundation models are proprietary, so users must share their data with model owners to fine-tune the models, which is costly and raise privacy concerns. Moreover, fine-tuning large foundation models is computation-intensive and impractical for most downstream users. In this paper, we propose Offsite-Tuning, a privacy-preserving and efficient transfer learning framework that can adapt billion-parameter foundation models to downstream data without access to the full model. In offsite-tuning, the model owner sends a light-weight adapter and a lossy compressed emulator to the data owner, who then fine-tunes the adapter on the downstream data with the emulator's assistance. The fine-tuned adapter is then returned to the model owner, who plugs it into the full model to create an adapted foundation model. Offsite-tuning preserves both parties' privacy and is computationally more efficient than the existing fine-tuning methods that require access to the full model weights. We demonstrate the effectiveness of offsite-tuning on various large language and vision foundation models. Offsite-tuning can achieve comparable accuracy as full model fine-tuning while being privacy-preserving and efficient, achieving 6.5x speedup and 5.6x memory reduction. Code is available at https://github.com/mit-han-lab/offsite-tuning.
TL;DR: In this paper , the authors examined the effect of increasing the number of model parameters on the performance of foundation models in downstream tasks such as rotated object detection and semantic segmentation, and proposed an effective method for scaling up and fine-tuning a vision transformer in the remote sensing field.
Abstract: As the potential of foundation models in visual tasks has garnered significant attention, pretraining these models before downstream tasks has become a crucial step. The three key factors in pretraining foundation models are the pretraining method, the size of the pretraining dataset, and the number of model parameters. Recently, research in the remote sensing field has focused primarily on the pretraining method and the size of the dataset, with limited emphasis on the number of model parameters. This paper addresses this gap by examining the effect of increasing the number of model parameters on the performance of foundation models in downstream tasks such as rotated object detection and semantic segmentation. We pretrained foundation models with varying numbers of parameters, including 86M, 605.26M, 1.3B, and 2.4B, to determine whether performance in downstream tasks improved with an increase in parameters. To the best of our knowledge, this is the first billion-scale foundation model in the remote sensing field. Furthermore, we propose an effective method for scaling up and fine-tuning a vision transformer in the remote sensing field. To evaluate general performance in downstream tasks, we employed the DOTA v2.0 and DIOR-R benchmark datasets for rotated object detection, and the Potsdam and LoveDA datasets for semantic segmentation. Experimental results demonstrated that, across all benchmark datasets and downstream tasks, the performance of the foundation models and data efficiency improved as the number of parameters increased. Moreover, our models achieve the state-of-the-art performance on several datasets including DIOR-R, Postdam, and LoveDA.
TL;DR: In this paper , the authors assess the future trends of the supply and demand relationship of water-related ecosystem services (WRESs) in the Asian water tower and its downstream area, which is closely related to the production and livelihoods of billions of people.
TL;DR: The results support a model in which PDF signaling negatively modulates EYA levels to regulate seasonal physiology, linking the circadian clock to the modulation of seasonal adaptations.
Abstract: Organisms adapt to seasonal changes in photoperiod and temperature to survive; however, the mechanisms by which these signals are integrated in the brain to alter seasonal biology are poorly understood. We previously reported that EYES ABSENT (EYA) shows higher levels in cold temperature or short photoperiod and promotes winter physiology in Drosophila. Nevertheless, how EYA senses seasonal cues is unclear. Pigment-dispersing factor (PDF) is a neuropeptide important for regulating circadian output rhythms. Interestingly, PDF has also been shown to regulate seasonality, suggesting that it may mediate the function of the circadian clock in modulating seasonal physiology. In this study, we investigated the role of EYA in mediating the function of PDF on seasonal biology. We observed that PDF abundance is lower on cold and short days as compared with warm and long days, contrary to what was previously observed for EYA. We observed that manipulating PDF signaling in eya+ fly brain neurons, where EYA and PDF receptor are co-expressed, modulates seasonal adaptations in daily activity rhythm and ovary development via EYA-dependent and EYA-independent mechanisms. At the molecular level, altering PDF signaling impacted EYA protein abundance. Specifically, we showed that protein kinase A (PKA), an effector of PDF signaling, phosphorylates EYA promoting its degradation, thus explaining the opposite responses of PDF and EYA abundance to changes in seasonal cues. In summary, our results support a model in which PDF signaling negatively modulates EYA levels to regulate seasonal physiology, linking the circadian clock to the modulation of seasonal adaptations.
TL;DR: In this article , the authors demonstrate 200 Gb/s downstream transmission for a passive optical network (PON) with a minimal coherent receiver at optical network unit (ONU), which employs heterodyne detection and requires a single photodiode, a 3-dB optical coupler and a local oscillator.
Abstract: We experimentally demonstrate 200 Gb/s/ $\lambda $ downstream transmission for a passive optical network (PON) with a minimal coherent receiver at optical network unit (ONU). The low complexity receiver required at the cost-sensitive ONU employs heterodyne detection and requires a single photodiode, a 3-dB optical coupler and a local oscillator, and is the simplest coherent receiver. At the optical line terminal (OLT), digital pre-emphasis and Alamouti-coding is applied to the 50 GBaud signal. A power budget of 29 dB is experimentally demonstrated, increasing to 34.5 dB if a balanced receiver is utilized.
TL;DR: Pre-training biases propagate to downstream tasks in text summarization, as shown in a case study.
Abstract: Faisal Ladhak, Esin Durmus, Mirac Suzgun, Tianyi Zhang, Dan Jurafsky, Kathleen McKeown, Tatsunori Hashimoto. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
TL;DR: In this article , the authors demonstrate the first real-time TFDMA coherent PON system with single-DAC and single-ADC ONUs, which can support up to 256 end users, and peak line rates of 100/200 Gb/s in the upstream/downstream, respectively.
Abstract: We demonstrate the first real-time TFDMA coherent PON system with single-DAC and single-ADC ONUs, which can support up to 256 end users, and peak line rates of 100/200 Gb/s in the upstream/downstream, respectively.