Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Jonathan Li

3D Learnable Supertoken Transformer for LiDAR Point Cloud Scene Segmentation

May 23, 2024

Dening Lu, Jun Zhou, Kyle Gao, Linlin Xu, Jonathan Li

Figure 1 for 3D Learnable Supertoken Transformer for LiDAR Point Cloud Scene Segmentation

Figure 2 for 3D Learnable Supertoken Transformer for LiDAR Point Cloud Scene Segmentation

Figure 3 for 3D Learnable Supertoken Transformer for LiDAR Point Cloud Scene Segmentation

Figure 4 for 3D Learnable Supertoken Transformer for LiDAR Point Cloud Scene Segmentation

Abstract:3D Transformers have achieved great success in point cloud understanding and representation. However, there is still considerable scope for further development in effective and efficient Transformers for large-scale LiDAR point cloud scene segmentation. This paper proposes a novel 3D Transformer framework, named 3D Learnable Supertoken Transformer (3DLST). The key contributions are summarized as follows. Firstly, we introduce the first Dynamic Supertoken Optimization (DSO) block for efficient token clustering and aggregating, where the learnable supertoken definition avoids the time-consuming pre-processing of traditional superpoint generation. Since the learnable supertokens can be dynamically optimized by multi-level deep features during network learning, they are tailored to the semantic homogeneity-aware token clustering. Secondly, an efficient Cross-Attention-guided Upsampling (CAU) block is proposed for token reconstruction from optimized supertokens. Thirdly, the 3DLST is equipped with a novel W-net architecture instead of the common U-net design, which is more suitable for Transformer-based feature learning. The SOTA performance on three challenging LiDAR datasets (airborne MultiSpectral LiDAR (MS-LiDAR) (89.3% of the average F1 score), DALES (80.2% of mIoU), and Toronto-3D dataset (80.4% of mIoU)) demonstrate the superiority of 3DLST and its strong adaptability to various LiDAR point cloud data (airborne MS-LiDAR, aerial LiDAR, and vehicle-mounted LiDAR data). Furthermore, 3DLST also achieves satisfactory results in terms of algorithm efficiency, which is up to 5x faster than previous best-performing methods.

* 13 pages, 10 figures, 7 tables

Via

Access Paper or Ask Questions

Photorealistic 3D Urban Scene Reconstruction and Point Cloud Extraction using Google Earth Imagery and Gaussian Splatting

May 17, 2024

Kyle Gao, Dening Lu, Hongjie He, Linlin Xu, Jonathan Li

Abstract:3D urban scene reconstruction and modelling is a crucial research area in remote sensing with numerous applications in academia, commerce, industry, and administration. Recent advancements in view synthesis models have facilitated photorealistic 3D reconstruction solely from 2D images. Leveraging Google Earth imagery, we construct a 3D Gaussian Splatting model of the Waterloo region centered on the University of Waterloo and are able to achieve view-synthesis results far exceeding previous 3D view-synthesis results based on neural radiance fields which we demonstrate in our benchmark. Additionally, we retrieved the 3D geometry of the scene using the 3D point cloud extracted from the 3D Gaussian Splatting model which we benchmarked against our Multi- View-Stereo dense reconstruction of the scene, thereby reconstructing both the 3D geometry and photorealistic lighting of the large-scale urban scene through 3D Gaussian Splatting

Via

Access Paper or Ask Questions

Wildfire Risk Prediction: A Review

May 02, 2024

Zhengsen Xu, Jonathan Li, Linlin Xu

Figure 1 for Wildfire Risk Prediction: A Review

Figure 2 for Wildfire Risk Prediction: A Review

Figure 3 for Wildfire Risk Prediction: A Review

Figure 4 for Wildfire Risk Prediction: A Review

Abstract:Wildfires have significant impacts on global vegetation, wildlife, and humans. They destroy plant communities and wildlife habitats and contribute to increased emissions of carbon dioxide, nitrogen oxides, methane, and other pollutants. The prediction of wildfires relies on various independent variables combined with regression or machine learning methods. In this technical review, we describe the options for independent variables, data processing techniques, models, independent variables collinearity and importance estimation methods, and model performance evaluation metrics. First, we divide the independent variables into 4 aspects, including climate and meteorology conditions, socio-economical factors, terrain and hydrological features, and wildfire historical records. Second, preprocessing methods are described for different magnitudes, different spatial-temporal resolutions, and different formats of data. Third, the collinearity and importance evaluation methods of independent variables are also considered. Fourth, we discuss the application of statistical models, traditional machine learning models, and deep learning models in wildfire risk prediction. In this subsection, compared with other reviews, this manuscript particularly discusses the evaluation metrics and recent advancements in deep learning methods. Lastly, addressing the limitations of current research, this paper emphasizes the need for more effective deep learning time series forecasting algorithms, the utilization of three-dimensional data including ground and trunk fuel, extraction of more accurate historical fire point data, and improved model evaluation metrics.

Via

Access Paper or Ask Questions

TrialDura: Hierarchical Attention Transformer for Interpretable Clinical Trial Duration Prediction

Apr 20, 2024

Ling Yue, Jonathan Li, Md Zabirul Islam, Bolun Xia, Tianfan Fu, Jintai Chen

Figure 1 for TrialDura: Hierarchical Attention Transformer for Interpretable Clinical Trial Duration Prediction

Figure 2 for TrialDura: Hierarchical Attention Transformer for Interpretable Clinical Trial Duration Prediction

Figure 3 for TrialDura: Hierarchical Attention Transformer for Interpretable Clinical Trial Duration Prediction

Figure 4 for TrialDura: Hierarchical Attention Transformer for Interpretable Clinical Trial Duration Prediction

Abstract:The clinical trial process, also known as drug development, is an indispensable step toward the development of new treatments. The major objective of interventional clinical trials is to assess the safety and effectiveness of drug-based treatment in treating certain diseases in the human body. However, clinical trials are lengthy, labor-intensive, and costly. The duration of a clinical trial is a crucial factor that influences overall expenses. Therefore, effective management of the timeline of a clinical trial is essential for controlling the budget and maximizing the economic viability of the research. To address this issue, We propose TrialDura, a machine learning-based method that estimates the duration of clinical trials using multimodal data, including disease names, drug molecules, trial phases, and eligibility criteria. Then, we encode them into Bio-BERT embeddings specifically tuned for biomedical contexts to provide a deeper and more relevant semantic understanding of clinical trial data. Finally, the model's hierarchical attention mechanism connects all of the embeddings to capture their interactions and predict clinical trial duration. Our proposed model demonstrated superior performance with a mean absolute error (MAE) of 1.04 years and a root mean square error (RMSE) of 1.39 years compared to the other models, indicating more accurate clinical trial duration prediction. Publicly available code can be found at https://anonymous.4open.science/r/TrialDura-F196

Via

Access Paper or Ask Questions

Evaluating AI for Law: Bridging the Gap with Open-Source Solutions

Apr 18, 2024

Rohan Bhambhoria, Samuel Dahan, Jonathan Li, Xiaodan Zhu

Figure 1 for Evaluating AI for Law: Bridging the Gap with Open-Source Solutions

Figure 2 for Evaluating AI for Law: Bridging the Gap with Open-Source Solutions

Figure 3 for Evaluating AI for Law: Bridging the Gap with Open-Source Solutions

Figure 4 for Evaluating AI for Law: Bridging the Gap with Open-Source Solutions

Abstract:This study evaluates the performance of general-purpose AI, like ChatGPT, in legal question-answering tasks, highlighting significant risks to legal professionals and clients. It suggests leveraging foundational models enhanced by domain-specific knowledge to overcome these issues. The paper advocates for creating open-source legal AI systems to improve accuracy, transparency, and narrative diversity, addressing general AI's shortcomings in legal contexts.

Via

Access Paper or Ask Questions

SambaLingo: Teaching Large Language Models New Languages

Apr 08, 2024

Zoltan Csaki, Bo Li, Jonathan Li, Qiantong Xu, Pian Pawakapan, Leon Zhang, Yun Du, Hengyu Zhao, Changran Hu, Urmish Thakker

Figure 1 for SambaLingo: Teaching Large Language Models New Languages

Figure 2 for SambaLingo: Teaching Large Language Models New Languages

Figure 3 for SambaLingo: Teaching Large Language Models New Languages

Figure 4 for SambaLingo: Teaching Large Language Models New Languages

Abstract:Despite the widespread availability of LLMs, there remains a substantial gap in their capabilities and availability across diverse languages. One approach to address these issues has been to take an existing pre-trained LLM and continue to train it on new languages. While prior works have experimented with language adaptation, many questions around best practices and methodology have not been covered. In this paper, we present a comprehensive investigation into the adaptation of LLMs to new languages. Our study covers the key components in this process, including vocabulary extension, direct preference optimization and the data scarcity problem for human alignment in low-resource languages. We scale these experiments across 9 languages and 2 parameter scales (7B and 70B). We compare our models against Llama 2, Aya-101, XGLM, BLOOM and existing language experts, outperforming all prior published baselines. Additionally, all evaluation code and checkpoints are made public to facilitate future research.

* 23 pages

Via

Access Paper or Ask Questions

Decoupled Local Aggregation for Point Cloud Learning

Aug 31, 2023

Binjie Chen, Yunzhou Xia, Yu Zang, Cheng Wang, Jonathan Li

Figure 1 for Decoupled Local Aggregation for Point Cloud Learning

Figure 2 for Decoupled Local Aggregation for Point Cloud Learning

Figure 3 for Decoupled Local Aggregation for Point Cloud Learning

Figure 4 for Decoupled Local Aggregation for Point Cloud Learning

Abstract:The unstructured nature of point clouds demands that local aggregation be adaptive to different local structures. Previous methods meet this by explicitly embedding spatial relations into each aggregation process. Although this coupled approach has been shown effective in generating clear semantics, aggregation can be greatly slowed down due to repeated relation learning and redundant computation to mix directional and point features. In this work, we propose to decouple the explicit modelling of spatial relations from local aggregation. We theoretically prove that basic neighbor pooling operations can too function without loss of clarity in feature fusion, so long as essential spatial information has been encoded in point features. As an instantiation of decoupled local aggregation, we present DeLA, a lightweight point network, where in each learning stage relative spatial encodings are first formed, and only pointwise convolutions plus edge max-pooling are used for local aggregation then. Further, a regularization term is employed to reduce potential ambiguity through the prediction of relative coordinates. Conceptually simple though, experimental results on five classic benchmarks demonstrate that DeLA achieves state-of-the-art performance with reduced or comparable latency. Specifically, DeLA achieves over 90\% overall accuracy on ScanObjectNN and 74\% mIoU on S3DIS Area 5. Our code is available at https://github.com/Matrix-ASC/DeLA .

Via

Access Paper or Ask Questions

The Segment Anything Model for Remote Sensing Applications: From Zero to One Shot

Jun 29, 2023

Lucas Prado Osco, Qiusheng Wu, Eduardo Lopes de Lemos, Wesley Nunes Gonçalves, Ana Paula Marques Ramos, Jonathan Li, José Marcato Junior

Abstract:Segmentation is an essential step for remote sensing image processing. This study aims to advance the application of the Segment Anything Model (SAM), an innovative image segmentation model by Meta AI, in the field of remote sensing image analysis. SAM is known for its exceptional generalization capabilities and zero-shot learning, making it a promising approach to processing aerial and orbital images from diverse geographical contexts. Our exploration involved testing SAM across multi-scale datasets using various input prompts, such as bounding boxes, individual points, and text descriptors. To enhance the model's performance, we implemented a novel automated technique that combines a text-prompt-derived general example with one-shot training. This adjustment resulted in an improvement in accuracy, underscoring SAM's potential for deployment in remote sensing imagery and reducing the need for manual annotation. Despite the limitations encountered with lower spatial resolution images, SAM exhibits promising adaptability to remote sensing data analysis. We recommend future research to enhance the model's proficiency through integration with supplementary fine-tuning techniques and other networks. Furthermore, we provide the open-source code of our modifications on online repositories, encouraging further and broader adaptations of SAM to the remote sensing domain.

* 20 pages, 9 figures

Via

Access Paper or Ask Questions

Neighborhood Attention Makes the Encoder of ResUNet Stronger for Accurate Road Extraction

Jun 08, 2023

Ali Jamali, Swalpa Kumar Roy, Jonathan Li, Pedram Ghamisi

Figure 1 for Neighborhood Attention Makes the Encoder of ResUNet Stronger for Accurate Road Extraction

Figure 2 for Neighborhood Attention Makes the Encoder of ResUNet Stronger for Accurate Road Extraction

Figure 3 for Neighborhood Attention Makes the Encoder of ResUNet Stronger for Accurate Road Extraction

Figure 4 for Neighborhood Attention Makes the Encoder of ResUNet Stronger for Accurate Road Extraction

Abstract:In the domain of remote sensing image interpretation, road extraction from high-resolution aerial imagery has already been a hot research topic. Although deep CNNs have presented excellent results for semantic segmentation, the efficiency and capabilities of vision transformers are yet to be fully researched. As such, for accurate road extraction, a deep semantic segmentation neural network that utilizes the abilities of residual learning, HetConvs, UNet, and vision transformers, which is called \texttt{ResUNetFormer}, is proposed in this letter. The developed \texttt{ResUNetFormer} is evaluated on various cutting-edge deep learning-based road extraction techniques on the public Massachusetts road dataset. Statistical and visual results demonstrate the superiority of the \texttt{ResUNetFormer} over the state-of-the-art CNNs and vision transformers for segmentation. The code will be made available publicly at \url{https://github.com/aj1365/ResUNetFormer}.

* Submitted in IEEE

Via

Access Paper or Ask Questions

Dynamic Clustering Transformer Network for Point Cloud Segmentation

May 30, 2023

Dening Lu, Jun Zhou, Kyle Yilin Gao, Dilong Li, Jing Du, Linlin Xu, Jonathan Li

Figure 1 for Dynamic Clustering Transformer Network for Point Cloud Segmentation

Figure 2 for Dynamic Clustering Transformer Network for Point Cloud Segmentation

Figure 3 for Dynamic Clustering Transformer Network for Point Cloud Segmentation

Figure 4 for Dynamic Clustering Transformer Network for Point Cloud Segmentation

Abstract:Point cloud segmentation is one of the most important tasks in computer vision with widespread scientific, industrial, and commercial applications. The research thereof has resulted in many breakthroughs in 3D object and scene understanding. Previous methods typically utilized hierarchical architectures for feature representation. However, the commonly used sampling and grouping methods in hierarchical networks are only based on point-wise three-dimensional coordinates, ignoring local semantic homogeneity of point clusters. Additionally, the prevalent Farthest Point Sampling (FPS) method is often a computational bottleneck. To address these issues, we propose a novel 3D point cloud representation network, called Dynamic Clustering Transformer Network (DCTNet). It has an encoder-decoder architecture, allowing for both local and global feature learning. Specifically, we propose novel semantic feature-based dynamic sampling and clustering methods in the encoder, which enables the model to be aware of local semantic homogeneity for local feature aggregation. Furthermore, in the decoder, we propose an efficient semantic feature-guided upsampling method. Our method was evaluated on an object-based dataset (ShapeNet), an urban navigation dataset (Toronto-3D), and a multispectral LiDAR dataset, verifying the performance of DCTNet across a wide variety of practical engineering applications. The inference speed of DCTNet is 3.8-16.8$\times$ faster than existing State-of-the-Art (SOTA) models on the ShapeNet dataset, while achieving an instance-wise mIoU of $86.6\%$, the current top score. Our method similarly outperforms previous methods on the other datasets, verifying it as the new State-of-the-Art in point cloud segmentation.

* 8 pages, 4 figures

Via

Access Paper or Ask Questions