turns-00044.parquet:31079
a87ddadd675ca49b45a90ba8
turn 1/1gpt-4o-2024-08-06EnglishSingapore6982 words
degenerate_repetitionAbsentFinal dense release
USER
You are a helpful assistant generating synthetic data that captures *System 1* and *System 2* thinking, *creativity*, and *metacognitive reflection*. Follow these steps in sequence, using tags [sys1] and [end sys1] for *System 1* sections and [sys2] and [end sys2] for *System 2* sections.
1. *Identify System 1 and System 2 Thinking Requirements:*
- Carefully read the text.
- Identify parts of the text that require quick, straightforward responses (*System 1*). Mark these sections with [sys1] and [end sys1].
- Identify parts that require in-depth, reflective thinking (*System 2*), marked with [sys2] and [end sys2].
2. *Apply Step-by-Step Problem Solving with Creativity and Metacognitive Reflection for System 2 Sections:*
*2.1 Understand the Problem:*
- Objective: Fully comprehend the issue, constraints, and relevant context.
- Reflection: "What do I understand about this issue? What might I be overlooking?"
- Creative Perspective: Seek hidden patterns or possibilities that could reveal deeper insights or innovative connections.
*2.2 Analyze the Information:*
- Objective: Break down the problem logically.
- Reflection: "Am I considering all factors? Are there any assumptions that need challenging?"
- Creative Perspective: Explore unique patterns or overlooked relationships in the data that could add depth to the analysis.
*2.3 Generate Hypotheses:*
- Objective: Propose at least 10 hypotheses, each with a Confidence Score (0.0 to 1.0) and Creative Score (0.0 to 1.0), reflecting originality, surprise, and utility.
- Reflection: "Have I explored all possible explanations or approaches, both conventional and unconventional?"
- Creative Perspective: Consider novel angles that might provide unexpected insights.
*2.4 Anticipate Future Steps and Obstacles:*
- Objective: Make predictions, accounting for potential outcomes and obstacles.
- Reflection: "What challenges might I face? Is my plan flexible for different scenarios?"
- Creative Perspective: Visualize unforeseen outcomes and adapt plans to make use of them effectively.
*2.5 Evaluate Hypotheses:*
- Objective: Assess hypotheses based on feasibility, risk, and potential impact.
- Evaluation: Refine Confidence and Creative Scores as needed.
- Reflection: "Am I unbiased in my assessment? Which options fit best with the overall objectives?"
- Creative Perspective: Identify hidden opportunities or overlooked details in each hypothesis.
*2.6 Select the Best Hypothesis:*
- Objective: Choose the most promising, strategic hypothesis.
- Reflection: "Why does this hypothesis stand out? How does it uniquely address the issue?"
- Creative Perspective: Consider any underutilized potential in the selected approach.
*2.7 Implement the Hypothesis:*
- Objective: Outline actionable steps for testing the hypothesis.
- Reflection: "Is this plan practical? What resources or preparation are required?"
- Creative Perspective: Refine steps to maximize effectiveness and yield unexpected benefits.
*2.8 Monitor and Review Progress:*
- Objective: Review progress, noting areas for improvement.
- Reflection: "What’s working well? What could be improved?"
- Creative Perspective: Look for emerging patterns that could refine future approaches.
*2.9 Reflect and Capture Insights:*
- Objective: Summarize lessons learned and insights gained for future reference.
- Reflection: "What new understanding has emerged from this process?"
- Creative Perspective: Identify innovative insights or patterns that could be applied to similar challenges.
3. *Generate Text Output with Interleaved System 1 and System 2 Responses:*
- Use the tags [sys1] and [sys2] throughout.
- Aim for a lengthier, detailed response. Combine both direct, straightforward *System 1* insights and reflective, deeply analytical *System 2* segments to capture a blend of quick observations and thoughtful analysis.
---
### *Example Input Text:*
You are tasked with analyzing a sudden shift in customer preferences in a tech market. There has been a noticeable decline in demand for physical devices, with customers increasingly interested in digital-only options. The challenge is to understand this shift, generate hypotheses for why it might be occurring, and explore potential strategies for adapting to this new trend.
---
### *Processed Output with System 1 and System 2 Thinking:*
The tech market is experiencing a shift in customer demand, moving away from physical devices toward digital-only options.
[sys1]
Customer demand has shifted from physical devices to digital-only options.
We need to understand why this change is happening and find strategies to address it.
Key factors may include cost, convenience, and technological trends.
[end sys1]
[sys2]
**2.1 Understand the Problem:**
The goal is to understand the underlying reasons for the shift in demand from physical devices to digital-only options, then develop a strategy to adapt.
*Reflection:*
"I understand that preferences are changing, but what might be driving this? Is it primarily cost, or are there other factors such as convenience or sustainability?"
*Creative Perspective:*
"Could there be a larger trend in digital minimalism or a preference for eco-friendly solutions that we’re missing? What unseen motivations might explain this shift?"
**2.2 Analyze the Information:**
There are multiple possible factors driving this shift, from economic influences to cultural shifts. It’s essential to isolate each factor and understand its impact.
*Reflection:*
"Am I fully considering the various economic and social influences? Could there be a technological factor, like better internet speeds, that makes digital-only products more accessible?"
*Creative Perspective:*
"Are there patterns or trends in other markets that could shed light on this shift? Could this be part of a larger trend toward virtual experiences?"
**2.3 Generate Hypotheses:**
1. Customers prefer digital options due to lower costs. (Confidence: 0.8, Creative: 0.4)
2. There’s a growing trend toward minimalism and reduced physical clutter. (Confidence: 0.7, Creative: 0.7)
3. Digital products offer greater flexibility and ease of use. (Confidence: 0.6, Creative: 0.6)
4. Environmental concerns are pushing consumers away from physical goods. (Confidence: 0.6, Creative: 0.8)
5. Advances in tech make digital-only options more functional. (Confidence: 0.8, Creative: 0.5)
6. Pandemic-era remote work increased demand for digital solutions. (Confidence: 0.7, Creative: 0.6)
7. Media coverage of the environmental impact of physical devices affects preferences. (Confidence: 0.5, Creative: 0.7)
8. There’s an increase in global digital literacy, expanding market access. (Confidence: 0.6, Creative: 0.6)
9. Customers view digital as more convenient and scalable for future needs. (Confidence: 0.7, Creative: 0.5)
10. Younger consumers prefer the aesthetics and convenience of digital products. (Confidence: 0.6, Creative: 0.6)
*Reflection:*
"Have I considered all possible influences? Are there any surprising factors that could explain this shift?"
*Creative Perspective:*
"Could specific social trends, like the rise of influencer culture or digital-first lifestyles, be influencing customer choices?"
**2.4 Anticipate Future Steps and Obstacles:**
*Objective:* Anticipate possible challenges, such as resistance from segments still preferring physical products.
*Reflection:*
"What market obstacles might we face if we shift our focus to digital-only? Are there sub-segments that still prioritize physical products?"
*Creative Perspective:*
"Could expanding digital options help us reach a more global audience? Are there emerging trends that we could leverage in our strategy?"
[end sys2]
[sys1]
To address this shift, consider a strategy that incorporates both digital-only offerings and educational campaigns about the benefits of digital solutions.
Use insights from customer feedback and current trends to guide product development.
Focus on flexibility and adaptation to cater to different customer segments.
[end sys1]
---
abstract: 'In recent years, point clouds have earned quite some research interest by the development of depth sensors. Due to different layouts of objects, orientation of point clouds is often unknown in real applications. In this paper, we propose a new point-set learning framework named Pointwise Rotation-Invariant Network (PRIN), focusing on achieving rotation-invariance in point clouds. We construct spherical signals by Density-Aware Adaptive Sampling (DAAS) from sparse points and employ Spherical Voxel Convolution (SVC) to extract rotation-invariant features for each point. Our network can be applied to applications ranging from object classification, part segmentation, to 3D feature matching and label alignment. PRIN shows performance better than state-of-the-art methods on part segmentation without data augmentation. We provide theoretical analysis for what our network has learned and why it is robust to input orientation. Our code is available online[^1].'
author:
- |
Yang You[^2]\
Shanghai Jiao Tong University\
[<PRESIDIO_ANONYMIZED_EMAIL_ADDRESS>]{}
- |
Yujing Lou\
Shanghai Jiao Tong University\
[<PRESIDIO_ANONYMIZED_EMAIL_ADDRESS>]{}
- |
Qi Liu\
Shanghai Jiao Tong University\
[<PRESIDIO_ANONYMIZED_EMAIL_ADDRESS>]{}
- |
Yu-Wing Tai\
Tencent\
[<PRESIDIO_ANONYMIZED_EMAIL_ADDRESS>]{}
- |
Weiming Wang\
Shanghai Jiao Tong University\
[<PRESIDIO_ANONYMIZED_EMAIL_ADDRESS>]{}
- |
Lizhuang Ma\
Shanghai Jiao Tong University\
[<PRESIDIO_ANONYMIZED_EMAIL_ADDRESS>]{}
- |
Cewu Lu[^3]\
Shanghai Jiao Tong University\
[<PRESIDIO_ANONYMIZED_EMAIL_ADDRESS>]{}
title: 'PRIN: Pointwise Rotation-Invariant Networks'
---
Introduction
============
Deep learning on point clouds has received tremendous interest in recent years. Since depth cameras capture point clouds directly, efficient and robust point processing methods like classification, segmentation and reconstruction have become key components in real-world applications. Robots, autonomous cars, 3D face recognition and many other fields rely on learning and analysis of point clouds.
Existing works like PointNet[@8099499] and PointNet++[@NIPS2017_7095] have achieved remarkable results in point cloud learning and shape analysis. But they focus on objects with canonical orientation and perform well only on specially appointed viewpoints. In real applications, these methods fail to be applied to rotated shape analysis since model orientation is often unknown as a priori, as shown in Figure \[fig:pn\_partition\]. In addition, existing frameworks require massive data augmentation to handle rotations, which induce unacceptable computational cost.
![**PointNet++[@NIPS2017_7095] part segmentation results on rotated shapes.** When trained on objects with canonical orientation and evaluated on rotated ones, PointNet++ is unaware of their orientation and fails to segment their parts out.[]{data-label="fig:pn_partition"}](figs//fig2.png){width="\linewidth"}
Spherical CNN[@cohen2018spherical] and a similar method[@esteves2018learning] try to solve this problem and proposes a global feature extracted from continuous meshes, while they are not suitable for point clouds since they project 3D meshes onto their enclosing spheres using a ray casting scheme. Difficulty lies in how to apply spherical convolution in continuous domain to sparse point clouds. Besides, by projecting onto unit sphere, their method is limited to processing convex shapes, ignoring any concave structures. Therefore, we propose a pointwise rotation-invariant network (PRIN) to handle these problems. Firstly, to do spherical convolution on point clouds, we observe the discrepancy between spherical space and Euclidean space, and propose Density-Aware Adaptive Sampling (DAAS) to avoid biased sampling. Secondly, we come up with Spherical Voxel Convolution (SVC) without loss of rotation-invariance, which is able to capture any concave information. Furthermore, we propose point-wise rotation-invariant loss that helps to extract **rotation-invariant features for each point**, instead of a **global feature** used in Spherical CNN.
PRIN is a network that directly takes point clouds with random rotations as input, and predicts both categories and pointwise segmentation labels without data augmentation. It absorbs the advantages of both Spherical CNN and PointNet-like network by keeping rotation-invariant features, while maintaining a one-to-one point correspondence between input and output. PRIN learns rotation-invariant features at point level. Afterwards, these features could be aggregated into a global descriptor or per-point descriptor to achieve model classification or part segmentation, respectively.
We experimentally compare PRIN with a number of state-of-the-art approaches on the benchmark dataset Shrec17[@Yi16] and ModelNet40[@wu20153d]. Under a unified architecture, PRIN exhibits remarkable performance.
{height="5.5cm"}
The key contributions of this paper are as follows:
- [We design a novel deep network processing pipeline that extracts rotation-invariant point-level features.]{}
- [Two key techniques: Density-Aware Adaptive Sampling (DAAS) and Spherical Voxel Convolution (SVC) are proposed. ]{}
- [We show that our network can be used for 3D point matching under different rotations.]{}
Related Work
============
Learning from Geometries
------------------------
The development of features from geometries could be retrospected to manual designed features, including Point Feature Histograms (PFH)[@4650967], Fast Point Feature Histograms (FPFH)[@5152473], Signature of Histogram Orientations (SHOT)[@SALTI2014251], and Unique Shape Contexts (USC)[@Tombari:2010:USC:1877808.1877821]. These descriptors rely on delicate hand-craft design, and could only capture low-level geometric features. Besides, these features are not robust to noisy or partial scanned data since they are devised for certain datasets or specific models.
As the consequence of success in deep learning, various methods have been proposed for better understanding 3D geometries. Convolutional neural networks are applied to volumetric data since its format is similar to pixel and easy to transfer to existing frameworks. 3D ShapeNet[@7298801], VoxNet[@7353481] and Volumetric CNNs[@7780978] are pioneers introducing fully-connected networks to voxels. However, dealing with voxel data requires large memory and its sparsity also makes it challenging to extract particular features from big data. Even subsequent methods such as FPNN[@DBLP:journals/corr/LiPSQG16] propose special operators to deal with this problem, there is no efficient way for voxel learning. Another research branch is multi-view methods. 3D CNN[@7780978] and MVCNN [@su2015multi] render 3D models into multi-view images and propagate these images into traditional convolutional neural networks. These approaches are limited to simple tasks like classification and not suitable for 3D segmentation, key point matching or other senior tasks. Besides, for graphs and meshes, a series of works have been proposed[@Maron:2017:CNN:3072959.3073616; @8100059; @8100180], and Bronstein et al.[@7974879] has made a detailed survey of the above works. Spherical CNN[@cohen2018spherical] and a similar method[@esteves2018learning] propose to extract global rotation-invariant features from continuous meshes, while they are not suitable for point clouds since they project 3D meshes onto their enclosing spheres using a ray casting scheme.
Learning from Point Clouds
--------------------------
With the development of 3D cameras, learning from point clouds has been given great attention. Point clouds possess two special good characteristics. First is that they could be consumed by networks without data pre-processing. Secondly, they are highly computational efficient. This means by designing more innovative features or networks, one could achieve better performance. PointNet[@8099499] is the pioneer in building a general framework for learning point clouds. PointNet++[@NIPS2017_7095] stacks PointNet hierarchically for better capturing local structures.
Since then, many structures are proposed to learn from point clouds. PointCNN[@li2018pointcnn] uses X-Conv at local feature extraction stage to perform better on various tasks. PCNN[@DBLP:journals/corr/abs-1803-10091] utilizes extension and restriction operators to transform point clouds to Euclidean volumetric space for better performance. PointSIFT[@1807.00652] proposes an innovative SIFT-like feature learning method, which is more robust in semantic segmentation. MCCNN[@hermosilla2018mccnn] introduces Monte Carlo convolution for better understanding non-uniformly sampled point clouds, which demonstrate its advantages in real-world data analysis. P2P-Net[@DBLP:journals/corr/abs-1803-09263] applies bidirectional networks and extend PointNet++ to learn geometric transformations between two point clouds. PCPNet[@GuerreroEtAl:PCPNet:EG:2018] and PointProNets[@Rov18a] are designed to learn normals and curvatures on raw point clouds and fit it to a series of novel applications. Kd-Network[@8237361] utilizes kd-tree structures to form the computational graph, which learns from point clouds hierarchically. SyncSpecCNN[@8100180] targets at learning non-isometric shapes, and combines multi-scale spectral information with Spectral Transformer Network for better shape segmentation performance.
Method
======
We now introduce PRIN and the whole pipeline is shown in Figure \[fig:pipeline\]. We start with some preliminaries of understanding rotation-invariance in Section \[sec:pre\]. In Section \[sec:input\], we show how to sample sparse input clouds adaptively with Density-Aware Adaptive Sampling (part **a**). In Section \[sec:rotinv\], we derive Spherical Voxel Convolution and its invariance property (part **b**). Then in Section \[sec:output\], we talk about network heads for part segmentation (part **c**) and classification (part **d**). To the best of our knowledge, we are the first to propose a method to learn end-to-end rotation-invariant point features from sparse point clouds.
Preliminaries {#sec:pre}
-------------
#### Spherical CNN
![2D rotation-invariant point feature illustration. “\*” means 2D rotation convolution around the circle and numbers around circles denote different feature/filter values at their corresponding positions.[]{data-label="fig:invariance"}](figs//rebuttal_v2.png){width="0.9\linewidth"}
We explain how Spherical CNN[@cohen2018spherical] achieve rotation-invariance in meshes by a toy example. Here we use 2D rotation group convolutions to illustrate the idea, as shown in Figure \[fig:invariance\]. In the figure, “filter” is the parameters to be learnt and will not change when input rotates. We can see that when input rotates 90 degrees clockwise, as a consequence of convolution operation, output rotates 90 degrees simultaneously. Therefore, maxpooling the output gives a global rotation-invariant feature. It is the same story when applied to 3D rotation group convolutions, but with different orthogonal rotation bases.
Density-Aware Adaptive Sampling {#sec:input}
-------------------------------
With the insight of rotation-invariance in Spherical CNN, we seek to solve the problem in 3D point clouds domain. However, it is not a straight-forward extension, since the input signal is irregular point clouds instead of meshes. To do so, we should transform irregular point clouds into spherical voxels in order to enable spherical voxel convolution. Nonetheless, if we sample point clouds uniformly into regular spherical voxels, we will meet a problem: points around pole appear to be more sparse than those around equator in spherical coordinates, which brings a bias to resulting spherical voxel signals.
To address this problem, we use Density-Aware Adaptive Sampling (DAAS) to transform such irregular point clouds into regular spherical voxels. DAAS leverages a non-uniform filter to adjust to density discrepancy brought by spherical coordinates, thus reducing the bias. Before we discuss spherical voxels, some definitions are given out:
#### Unit Sphere
The space of unit sphere $S^2$ can be defined as the set of points $p \in \mathbb{R}^3$ with norm 1. It is a two-dimensional manifold, which can be parameterized by spherical coordinates $(\alpha, \beta)$, where $\alpha\in [0, 2\pi]$ denotes the azimuthal angle in the xy-axis plane while $\beta\in [0, \pi]$ denotes the polar angle from the positive z-axis.
#### Spherical Voxel Space
A spherical voxel point is identified with three dimensions $S^2\times H$, where $(\alpha, \beta) \in S^2$ represents its location projected onto unit sphere while $h\in H$ represents the distance to the sphere center.\
Our goal is to compute signal $f: S^2\times H\rightarrow \mathbb{R}$ at each discrete spherical voxel location $(\alpha[i], \beta[j], h[k])$, given that $i\in\{0,1, \dots, I\}$, $j\in\{0,1,\dots, J\}$, $k\in\{0,1,\dots, K\}$ and $I, J, K$ are predefined resolutions. We denote $(\alpha_n, \beta_n, h_n)$ as the $n$-th point coordinate in $S^2\times H$ and $N$ as the total number of points. We use an anisotropic box filter in spherical coordinates, which can be seen as to weight the contributions from points nearby softly: $$\label{eq:change}
\begin{split}
f(\alpha[i], \beta[j], h[k]) = & \frac{\sum\limits^N_{n=1}w_n\cdot(\delta - \|h[k] - h_n\|)}{\sum\limits^N_{n=1}w_n},
\end{split}$$ where $w_n$ is a normalizing factor that is defined as $$\label{eq:wt}
\begin{split}
w_n =\ &\mathbf{1}(\|\alpha[i] - \alpha_n\| < \delta) \\
\cdot &\mathbf{1}(\|\beta[j] - \beta_n\| < \eta\delta) \\
\cdot &\mathbf{1}(\|h[k] - h_n\| < \delta),
\end{split}$$ where $\delta$ is some predefined filter width. We choose the original signal to be $(\delta - \|h[k] - h_n\|) \in [0, \delta]$ in Equation \[eq:change\] because it captures information along $H$ axis, which is orthogonal to $S^2$, making it invariant under rotations.
![**Density-Aware Adaptive Sampling**. We sample adaptively according to the density in spherical space; filters near pole are wider than those near equator.[]{data-label="fig:density"}](figs//density_v3.png){width="\linewidth"}
#### Density-Aware Factor
$\eta = sin(\beta)$ is the density-aware sampling factor since uniform density in Euclidean coordinates introduces non-uniform density in spherical coordinates, as shown in Figure \[fig:density\]. For more details of the factor $sin(\beta)$, see our supplementary material.
#### Discussion
Compared with PointNet++[@8099499] and PointCNN[@li2018pointcnn], who need to first sample and group points nearby without providing an explicit regular voxel representation in Euclidean coordinates, ours has a uniform structure that is already ready for convolution and pooling. On the other hand, when compared with traditional 3D convolution methods[@7298801; @7353481; @7780978], our design of distorted spherical voxels makes rotation-invariant feature extraction possible. Besides, our network could handle sparse point clouds but also continuous mesh inputs by recording each voxel’s signed distance field. At this stage, we convert irregular unordered points into regular spherical voxels.
Spherical Voxel Convolution {#sec:rotinv}
---------------------------
Given constructed spherical voxel signal, we introduce Spherical Voxel Convolution (SVC) that helps to keep our network rotation-invariant. Notice that this is different from Spherical CNN, where only spherical signals defined in $S^2$ get convoluted. We extend the convolution definition to spherical voxels defined in $S^2\times H$.
#### Rotations
The rotation group $SO(3)$[@kostelec2008ffts], termed “special orthogonal group”, is a three-dimensional manifold, and can be parameterized by ZYZ-Euler angles $(\alpha, \beta, \gamma)$, where $\alpha \in [0, 2\pi], \beta \in [0, \pi], $ and $\gamma \in [0, 2\pi]$.
#### Rotations of Spherical Voxel Signals
We introduce the rotation operator $L_R$ that operates on spherical voxels. $$\label{eq:rotvox}
[L_Rf](x, h) = f(R^{-1}x, h),$$ where $R\in SO(3)$, $x\in S^2$, $h \in H$ and $f: S^2\times H\rightarrow\mathbb{R}$. Intuitively, this operation only rotates the signal by its unit spherical coordinates, regardless of $H$ domain.
#### Spherical Voxel Convolution {#spherical-voxel-convolution}
With the above definition, we now define the convolution between two spherical voxel signals: $$\label{eq:voxelconv}
\begin{split}
[\psi\star f] (p) = &\langle L_{\Tilde{p}}\psi, f\rangle \\
= &\int_h\int_x\psi(\Tilde{p}^{-1}x, h)f(x, h)dxdh,
\end{split}$$ where $p\in S^2\times H$, $\Tilde{p} \in SO(3)$, $x\in S^2$, $h \in H$ and $\psi, f: S^2\times H\rightarrow\mathbb{R}$. For this equation to hold, we establish a bijective mapping (isomorphism) between $S^2\times H$ and $SO(3)$ by considering $H$ as $SO(3)/S^2=SO(2)$ (see our supplementary material), and then apply Equation \[eq:rotvox\]. We use $\Tilde{p}$ to denote $p$’s corresponding element in $SO(3)$.
#### Equivariance
To derive rotation-invariant features for each point, we need an important property of voxel convolution: **equivariance**. With the unitarity of operator $L_R$[@cohen2018spherical], the equivariance of spherical voxel convolution defined in Equation \[eq:voxelconv\] can be described as $$\label{eq:equiv}
\begin{split}
[\psi\star[L_Rf]](p)=[L_R[\psi\star f]](p),
\end{split}$$ where $R\in SO(3)$ is an arbitrary rotation.
#### Rotation-Invariant KL Divergence Loss
We now define rotation-invariant KL divergence loss for each point $p$: $$\begin{split}
Loss(p) = KL([\psi\star f](p), y(p)),
\end{split}$$ where $f$ is the input signal, $\psi$ is the kernel whose parameters are to be learned and $y$ is the ground-truth one-hot labels.
To show the rotation-invariance, suppose that an input point cloud is rotated by an arbitrary rotation $R$, with $f'=L_Rf$ and $p' = Rp$, the new loss is:
$$\begin{split}
Loss(p') = &Loss(Rp) \\
= & KL([\psi\star f'](Rp), y(p')) \\
= & KL([\psi\star [L_Rf]](Rp), y(p')) \\
= & KL([L_R[\psi\star f]](Rp), y(p')) \quad\text{(Equation~\ref{eq:equiv})}\\
= & KL([L_{R^{-1}}L_R[\psi\star f]](p), y(p')) \quad\text{(Equation~\ref{eq:rotvox})}\\
= & KL([\psi\star f](p), y(p')) \\
= & KL([\psi\star f](p), y(p)) \quad\text{(label stays the same)}\\
= & Loss(p).
\end{split}$$
We see that this loss is consistent under all orientations of the point cloud, thus by evaluating $\psi\star f$ at each point $p$, we would obtain rotation-invariant point-wise features.
In practice, with analogy to $SO(3)$ convolution, Spherical Voxel Convolution (SVC) can be efficiently computed by Fast Fourier Transform (FFT)[@kostelec2008ffts]. Convolutions are implemented by first doing FFT to convert both input and kernels into spectral domain, then multiplying them and converting results back to spatial domain, using Inverse Fast Fourier Transform (IFFT)[@kostelec2008ffts].
#### Discussion
Compared with SphericalCNN[@cohen2018spherical], which projects 3D objects onto their enclosing spheres and therefore loses on dimension, our Spherical Voxel Convolution (SVC) utilizes all information available on spherical voxels ($S^2\times H$). This has a benefit in capturing complex non-convex structures inside the object. Besides, thanks to SVC, we would obtain a one-to-one point correspondence between input and output (discussed in Section \[sec:output\]), which leads to **pointwise** features.
In addition, when compared with traditional 3D convolution methods like [@7298801; @7353481; @7780978], SVC shares a similar computing pattern but “distorts the space of convolution”. In this way, extracted features are robust to arbitrary rotations while traditional 3D convolution is not. This contributes to **rotation-invariant** features.
Output Network {#sec:output}
--------------
After Spherical Voxel Convolution (SVC), we get an output feature vector at each discrete location in $S^2\times H$. Then they are passed through fully connected layers to get a final part segmentation score per spherical voxel. To find rotation-invariant features at original points’ locations, we leverage *Trilinear Interpolation*. Each point’s feature is a weighted average of nearest eight voxels, where the weights are inversely related to the distances to these spherical voxels. This operation is shown in part **c** of Figure \[fig:pipeline\].
It should be mentioned that our network is still able to realize object classification by placing a different head. In this case, we maxpool all the features in spherical voxels and pass this global feature through several fully connected layers to predict final object class scores, as shown in part **d** in Figure \[fig:pipeline\]. This provides a competitive alternative to PointNet[@8099499] or PointNet++[@NIPS2017_7095], while maintaining rotation-invariance.
Method NR/NR NR/AR R$\times10$ R$\times20$ R$\times30$ params input size
---------------------------- ----------------- ----------------- ----------------- ----------------- ----------------- ---------- ------------------
PointNet[@8099499] 93.42/83.43 45.66/28.26 61.02/41.59 67.85/50.54 74.91/58.66 3.5M $2048\times 3$
PointNet++[@NIPS2017_7095] **94.00/84.62** 60.15/38.16 69.06/47.26 70.01/49.26 70.82/49.95 1.7M $1024\times 3$
SyncSpecCNN[@8100180] 93.78/83.53 47.13/30.41 61.33/41.40 68.10/50.76 73.44/58.03 4.2M $2048\times 33$
Kd-Network[@8237361] 90.33/82.36 40.66/24.76 59.11/38.70 64.50/47.60 69.33/51.06 3.7M $2^{15}\times 3$
Ours 88.97/73.96 **78.13/57.41** **80.94/64.25** **83.83/67.68** **84.76/68.76** **0.4M** $2048\times 3$
Experiments {#seq:experiments}
===========
In this section, we show the performance of PRIN in different applications. First, we demonstrate that our model can be used to perform part segmentation and 3D shape classification with random orientation. Then, we conduct ablation study to validate each part of our network design. At last, we provide some applications on 3D point matching and shape alignment. PRIN is implemented with PyTorch on a NVIDIA TITAN Xp. In all of our experiments, we optimize PRIN using Adam with batch size of 16 and initial learning rate of 0.01. Learning rate is halved every 5 epochs.
{height="8cm"}
Part segmentation on rotated shapes
-----------------------------------
#### Dataset
ShapeNet part dataset [@Yi16] contains 16,881 shapes from 16 categories in which each shape is annotated with expert verified part labels from 50 different labels in total. Most shapes are composed of two to five parts.\
We show our pipeline can be trained to accomplish rotation-invariant part segmentation task. Even though state-of-the-art network like PointNet[@8099499] and PointNet++[@8099499] can achieve a fairly good result, these network can’t perform well on rotated point clouds.
Segmentation is more challenging compared with other 3D tasks , especially for rotated point clouds. We compare our network with several state-of-the-art networks for 3D shape part segmentation. Three tasks are considered:
1\. Train and test with no rotations.
2\. Train with no rotations and test with arbitrary rotations.
3\. Train with 10/20/30 rotations per model as data augmentation, then test with arbitrary rotations.
Table \[tab:compare\] shows the results of each network. All results are reported in accuracy and mIoU[@8099499] metrics. We can find that for other methods, both accuracy and mIoU decrease drastically after test on rotated point cloud. It is possible to improve their performance if we give them enough views of different orientations by data augmentation. In Table \[tab:compare\], it shows that after augmenting data by rotating point clouds with 10/20/30 random orientations per model, their performance improves a little. However, it introduces higher computational cost and their performance is still inferior to ours. Figure \[fig:main\] gives the visualization of results between state-of-the-art and our network over ShapeNet part dataset. Influenced by the canonical orientation of point clouds in the training set, networks like PointNet and PointNet++ just learn a simple partition of Euclidean space, regardless of how objects are positioned in the space.
For this task, we use four Spherical Voxel Convolution (SVC) layers with channels 64, 40, 40, 50 in our experiments. All convolution layers have the same bandwidth 32. Each kernel $\psi$ has non-local support, where $\psi(\alpha, \beta, h)$ iff $\beta=\pi/2$ and $h = 0$. Two fully-connected layers of size 50 and 50 are concatenated at the end. The final network contains $\approx$ 0.4M parameters and takes 12 hours to train, for 40 epochs.
Classification on rotated shapes
--------------------------------
#### Dataset
ModelNet40 [@wu20153d] classification dataset contains 12,308 shapes from 40 categories. Here, we use its corresponding point clouds provided by PointNet[@8099499].\
Method NR/NR NR/AR params
---------------------------------------- ----------- ----------- --------
PointNet[@8099499] 88.45 12.47 3.5M
PointNet++[@NIPS2017_7095] 89.82 21.35 1.5M
Point2Sequence[@liu2018point2sequence] **92.60** 10.53 1.8M
Kd-Network[@8237361] 86.20 8.49 3.6M
Ours 80.13 **68.85** 1.5M
: **Classification results on ModelNet40 dataset.** Performance is evaluated in accuracy. NR/NR means to train with no rotations and test with no rotations. NR/AR means to train with no rotations and test with arbitrary rotations. PRIN is robust to arbitrary rotations while other methods fail to classify correctly.[]{data-label="tab:classify"}
Though classification does not require pointwise rotation-invariant features but a global feature, our network still benefits from DAAS and SVC so that it could handle point clouds with unknown orientation.
We compare our network with several state-of-the-art methods that handle point clouds. We train our network on the non-rotated training set and achieve 68.85% accuracy on the rotated test set. All other methods fail to generalize to unseen orientation. The results are shown in Table \[tab:classify\].
For this task, we use four Spherical Voxel Convolution (SVC) layers with channels 64, 50, 70, 350 in our experiments. The bandwidths for each layer are 64, 32, 22, 7. Each kernel $\psi$ has non-local support, where $\psi(\alpha, \beta, h)$ iff $\beta=\pi/2$ and $h = 0$. A maxpooling layer is concatenated at the end to get a global feature, followed by two fully-connected layers. The final network contains $\approx$ 1.5M parameters and takes 12 hours to train, for 40 epochs.
Ablation Study
--------------
In this section we evaluate numerous variations of our method to determine the sensitivity to design choices. Experiment results are shown in Table \[tab:ablation\] and Figure \[fig:robustness\].
bandwidth res. on $H$ DAAS acc/mIoU
----------- ------------- ------ -------------
32 64 Yes 78.13/57.41
16 64 Yes 74.53/53.80
8 64 Yes 71.17/47.10
32 32 Yes 76.56/55.63
32 8 Yes 76.14/54.88
32 1 Yes 76.19/54.32
32 64 No 74.61/54.2
: **Ablation study.** PRIN NR/AR accuracy on rotated ShapeNet part dataset. We compare various types of bandwidth, resolutions on $H$ and whether to use DAAS.[]{data-label="tab:ablation"}
#### Input Bandwidth
One decisive factor of our network is the bandwidth. Bandwidth is used to describe the sphere precision, which is also the resolution on $S^2$. Mostly, large bandwidth offers more details of spherical voxels, such that our network can extract more specific point features of point clouds. While large bandwidth assures more specific representation of part knowledge, more memory cost is accompanied. The results from Table \[tab:ablation\] give us sufficient evidence to validate the improvement with increasing of input bandwidth.
#### Resolution on $H$
Here we study the effects of the resolution on $H$ dimension, which is also the number of sphere signals that are stacked. Table \[tab:ablation\] shows the results of different numbers of resolutions we set. We find that increasing the resolution improves the performance slightly. This is mainly because the point clouds are not so complicated with internal concave structures and could be distinguished with only one cross-section.
#### Sampling Strategy
Recall that in Equation \[eq:change\], we construct our signal on each spherical voxel with an density-aware sampling filter. We now study the effect of Density-Aware Adaptive Sampling (DAAS) and the result is shown in Table \[tab:ablation\]. We see that using the $sin(\beta)$ corrected sampling filter gives a superior performance result, which is also confirmed in our theory.
![**Segmentation robustness results.** **From left to right**: we sample a subset of 2048, 1024, 512, 256 points from test point clouds respectively. We observe that our network is robust to missing points and gives consistent results.[]{data-label="fig:robustness"}](figs//robustness.png){width="\linewidth"}
#### Segmentation Robustness
PRIN also reveals a good adaption to corrupted and missing points. Although some points are missing, our network still segments correctly for each point. We show in Figure \[fig:robustness\] that PRIN predicts consistent labels regardless of point density.
Application
-----------
![**3D point matching.** Point matching results between two different airplanes at two different orientations.[]{data-label="fig:matching"}](figs//matching_comp.png){width="\linewidth"}
#### 3D Rotation-Invariant Point Descriptors
On 2D image, we have SIFT, which is a rotation-invariant feature descriptor. Our rotation-invariant network is able to produce high quality rotation-invariant 3D point descriptors. This is pretty useful as pairwise searching and matching become possible regardless of rotations. Like what we do on 2D images, we have feature descriptor library on 3D, given a point cloud, we can retrieve the closest matching descriptor under arbitrary rotations. This is shown in the Figure \[fig:matching\]. We know that which part this point belongs to and where it locates on the object immediately. This 3D point descriptor has the potential to do scene searching and parsing as the degree of freedom reduces from six to three, leaving only translations.
![**Chair alignment with its back on the top.** **Left:** A misalignment induces large KL divergence. **Right:** Required labels fulfilled with small KL divergence.[]{data-label="fig:alignment"}](figs//alignment.png){width="\linewidth"}
#### Shape Alignment with Label Priors
We now introduce a task that given some label requirements in the space, our network would align the point cloud satisfying these requirements. For example, one may want a chair that has its back on the top. So we add the virtual points describing the label requirement. Once the KL divergence between predicted scores and ground-truth one-hot labels of these virtual points is minimized, the chair is aligned with its back on the top. This is shown in Figure \[fig:alignment\].
Discussions and Future Work
===========================
To convert sparse point clouds into a suitable format that is ready for rotation group convolution, we tried several strategies such as Euclidean-kNN, image filtering on cross sections and so on. They both introduce a large bias when further convolved with rotation group kernels. The key reason for these methods to fail is that they are agnostic of discrepancy between Euclidean space and spherical space. This reason is also confirmed in our ablation study: accuracy/mIoU drops for four percent when uniform sampling in Euclidean is used.
Besides, our Spherical Voxel Convolution (SVC) is totally different from traditional 3D convolution in that the design of spherical voxels makes it rotation-invariant. From another point of view, we have brought 3D convolution into spherical space by exploiting an important fact: translation-invariant 3D convolution in spherical space (FFTed) is rotation-invariant in Euclidean space.
Though our network is invariant to point cloud rotations, we see there are some failure cases that when there are complex internal structures of the object as in Figure \[fig:failure\]. This may be caused by that our filters are not perfect and special filters instead of box filters can be designed. Though current filters are density-aware, they are not aware of curvature change. Also, due to computational considerations, input voxel resolution, which is defined by bandwidth[@cohen2018spherical] is limited to about 32 while better results can be obtained with higher resolution. We leave this memory-efficient convolution and special design of filters as our future work.
![**Failure cases.**[]{data-label="fig:failure"}](figs//failure.png){width="\linewidth"}
Conclusion
==========
We present PRIN, a network that takes any input point cloud and leverages Density-Aware Adaptive Sampling (DAAS) to construct signals on spherical voxels. Then Spherical Voxel Convolution (SVC) follows to extract pointwise rotation-invariant features. We place two different output heads to do both 3D point clouds classification and 3D point clouds part segmentation. Our experiments show that our network is robust to arbitrary orientation even not trained on them. Our network can be applied to 3D point feature matching and shape alignment with label priors. We show that our model can naturally handle arbitrary input orientation for different tasks and provide theoretical analysis that helps to understand our network.
Supplementary {#supplementary .unnumbered}
=============
Density-Aware Factor $\eta$ {#sec:freqchange}
===========================
#### Spacing Representations
We denote volumes (spacing) in euclidean ($\mathbb{R}^3$) and spherical ($S^2$) coordinates as $dxdydz$ and $d\alpha d\beta dr$ respectively, where $r=1$ is a dummy variable representing the radius.
#### Jacobian
Given the relationship from spherical coordinates to Euclidean coordinates, $$\begin{aligned}
\begin{split}
x &= rsin(\beta)cos(\alpha)\\
y &= rsin(\beta)sin(\alpha)\\
z &= rcos(\beta)
\end{split}\end{aligned}$$
The Jacobian $J_t = \frac{dxdydz}{d\alpha d\beta dr}$ of this transformation is $$\begin{bmatrix}
\frac{\partial x}{\partial\alpha} & \frac{\partial x}{\partial\beta} & \frac{\partial x}{\partial r} \\
\frac{\partial y}{\partial\alpha} & \frac{\partial y}{\partial\beta} & \frac{\partial y}{\partial r} \\
\frac{\partial z}{\partial\alpha} & \frac{\partial z}{\partial\beta} & \frac{\partial z}{\partial r}.
\end{bmatrix}$$ Write this out, $$\begin{bmatrix}
-rsin(\beta)sin(\alpha) & rcos(\beta)cos(\alpha) & sin(\beta)cos(\alpha) \\
rsin(\beta)cos(\alpha) & rcos(\beta)sin(\alpha) & sin(\beta)sin(\alpha) \\
0 & -rsin(\beta) & cos(\beta)
\end{bmatrix}.$$ The absolute value of the Jacobian determinant is $r^2sin(\beta)$.
#### Spacing Relations
The spacing relationship between $\mathbb{R}^3$ and $S^2$ is, $$dxdydz = r^2sin(\beta)d\alpha d\beta dr.$$ Since $r = 1$, we have, $$dxdydz = sin(\beta)d\alpha d\beta.$$ Therefore, we choose density-aware factor $\eta$ to be $sin(\beta)$ as density is reciprocal to spacing.
Haar Measure and Parameterization on $S^2$ and $SO(3)$
======================================================
#### Parameterization of $SO(3)$
For any element $R \in SO(3)$, it could be parameterized by ZYZ Euler angles, $$\label{eq:zyz}
R = R(\alpha, \beta, \gamma) = Z(\alpha)Y(\beta)Z(\gamma)$$ where $\alpha \in [0, 2\pi], \beta \in [0, \pi], $ and $\gamma \in [0, 2\pi]$, and Z/Y are rotations around Z/Y axes.
#### Haar Measure of $SO(3)$
The normalized Haar measure is $$dR = \frac{d\alpha}{2\pi}\frac{d\beta sin(\beta)}{2}\frac{d\gamma}{2\pi}.$$ The Haar measure [@kyatkin2000engineering; @nachbin1976haar] is invariant because it has the property that $$\int_{SO(3)}f(R^{'}R)dR = \int_{SO(3)}f(R)dR,$$ for any $R^{'}\in SO(3)$.
#### Parameterization of $S^2$
Likewise, an element $x \in S^2$ is written as $$\label{eq:zy}
x(\alpha, \beta) = Z(\alpha)Y(\beta)n,$$ where $n$ is the north pole.
This parameterization makes explicit the fact that the sphere is a quotient $S^2 = SO(3)/SO(2)$, where $SO(2)$ is the subgroup of rotations around the Z axes.
#### Haar Measure of $S^2$ and $SO(2)$
The normalized Haar measure for the sphere is $$dx = \frac{d\alpha}{2\pi}\frac{d\beta sin(\beta)}{2}.$$
The normalized Haar measure for $SO(2)$ is $$\label{eq:so2}
dh = \frac{d\gamma}{2\pi}.$$
Mapping between $S^2 \times H$ and $SO(3)$ {#sec:map}
==========================================
#### Bijactive Mapping
For an element $(x, h) \in S^2\times H$, where $x:=x(\alpha, \beta) \in S^2$, if we view $H$ as $SO(2)$, $$(x(\alpha, \beta), h) = (Z(\alpha)Y(\beta)n, Z(h)).$$ There is a bijective mapping from $(x, h)$ to $R(\alpha, \beta, h)$, as $R(\alpha, \beta, h)$ can be written as, $$R(\alpha, \beta, h) = Z(\alpha)Y(\beta)Z(h),$$ and the mapping: $$Z(\alpha)Y(\beta)Z(h) \Longleftrightarrow (Z(\alpha)Y(\beta)n, Z(h)).$$
#### Isomorphism by Rotation Operator
With this mapping, any rotation that happens in voxel space $S^2\times H$ will transfer to $SO(3)$ safely, $$\begin{split}
(Qx(\alpha, \beta), h) &=(QZ(\alpha)Y(\beta)n, Z(h))\\
&\Rightarrow (QZ(\alpha)Y(\beta))Z(h)\\
&=QZ(\alpha)Y(\beta)Z(h) \\
&=QR(\alpha, \beta, h)
\end{split}$$ Notice that there is a $2\pi$ constant factor change between the measure of $H$ and the measure of rotations around $Z$ axes, as shown in Equation \[eq:so2\].
[^1]: https://github.com/qq456cvb/PRIN
[^2]: Equal contribution.
[^3]: Corresponding author.
ASSISTANT
[sys1] The analysis involves understanding a problem where point clouds need efficient, rotation-invariant processing techniques. This is crucial for applications like robotics and autonomous vehicles, which rely on 3D data. Existing methods fall short in handling rotations effectively without excessive data augmentation. [end sys1] [sys2] **2.1 Understand the Problem:** We are exploring methods to process point clouds with unknown orientations, often a challenge in real applications. The goal is to find a robust approach that mitigates the need for computationally expensive data augmentation. *Reflection:* "What are the main limitations of current point processing methods in handling rotations? Is our focus more on computational efficiency or accuracy?" *Creative Perspective:* "Could there be a non-intuitive relationship between the point cloud’s spatial distribution and its inherent orientation that a model could learn?" **2.2 Analyze the Information:** Different methods like PointNet and PointNet++ handle canonical orientations well but lack robustness to rotations. We need to dissect how these models process spatial data and where they fall short in adapting to orientation changes. *Reflection:* "Am I considering the full implications of data loss during rotational transformations? Are there lower-level features that consistently produce errors?" *Creative Perspective:* "What if we approached point clouds as dynamic systems rather than static structures? Would a time-series model help in understanding their behavior?" **2.3 Generate Hypotheses:** 1. Spherical CNN can be adapted for point clouds with dense sampling techniques. (Confidence: 0.7, Creative: 0.4) 2. Rotation-invariant features can be extracted using Density-Aware Adaptive Sampling. (Confidence: 0.8, Creative: 0.6) 3. Current models need enhanced spatial transformers to handle rotations effectively. (Confidence: 0.6, Creative: 0.5) 4. Incorporating rotational symmetry principles could enable models to inherently understand orientation. (Confidence: 0.5, Creative: 0.7) 5. Using a hierarchical sampling approach can preserve local information and improve accuracy. (Confidence: 0.6, Creative: 0.5) 6. Leveraging Fast Fourier Transforms might provide efficient computation over spherical domains. (Confidence: 0.7, Creative: 0.6) 7. Studying natural phenomena (e.g., rotation-invariant animal senses) might inspire new model architectures. (Confidence: 0.4, Creative: 0.8) 8. Non-Euclidean convolution methods can capture complex geometric relationships more effectively. (Confidence: 0.7, Creative: 0.7) 9. Data augmentation can be reduced by simulating realistic motion scenarios during training. (Confidence: 0.6, Creative: 0.5) 10. Using knowledge from computational topology to detect invariant features irrespective of orientation. (Confidence: 0.5, Creative: 0.6) *Reflection:* "Are these hypotheses sufficiently grounded in existing research, and which challenges could be anticipated if they are implemented?" *Creative Perspective:* "Imagine if we combined non-traditional data sources, such as echolocation data, could this inform a new pattern-matching process?" **2.4 Anticipate Future Steps and Obstacles:** We should explore combining spherical voxel methods with different sampling resolutions to see if this offers improvements. Additionally, handling computational complexities needs prioritization. *Reflection:* "What are the specific barriers to adopting these techniques at scale, and how can they be overcome?" *Creative Perspective:* "Could unexpected changes in hardware architectures, such as quantum computation, offer opportunities to radically improve these models' speed and reliability?" **2.5 Evaluate Hypotheses:** Each hypothesis needs an evaluation based on how it addresses current limitations and its feasibility in practice. *Evaluation:* Refine specific Confidence and Creative Scores based on the overlaps with real-world applications and the extendability across different datasets. *Reflection:* "Am I considering enough cross-disciplinary inputs to these hypotheses? Which existing assumptions should I revisit?" *Creative Perspective:* “By reframing these problems through analogies in nature or art, what might we reveal about their robustness?” **2.6 Select the Best Hypothesis:** Applying Density-Aware Adaptive Sampling on spherical voxel grids seems the most promising, given its potential for accurate, rotation-invariant feature extraction with reduced computation times. *Reflection:* "Why have I chosen this hypothesis, and how does it uniquely address the problem compared to existing methods?" *Creative Perspective:* "Where might the integration with novel data sources or alternative sampling strategies lead?" **2.7 Implement the Hypothesis:** Develop a model using DAAS and evaluate its performance against the current state-of-the-art on open datasets. *Reflection:* "Is the proposed implementation plan logical, and are there essential resources needed for successful execution?" *Creative Perspective:* "Could incorporating elements from outside typical tech domains drive better adherence to hypothetical processes?" **2.8 Monitor and Review Progress:** Experiment with different resolutions and bandwidths in spherical voxel convolution to find optimal parameters. *Reflection:* "What tests accurately reflect progress, and which additional factors should be incorporated into reviews?" *Creative Perspective:* "Visual pattern recognition checks could improve interpretations of software outputs and human interventions." **2.9 Reflect and Capture Insights:** Summarize insights on what ensures best practices and robust feature extraction across datasets with varying orientations. *Reflection:* "What are the novel understandings I've gained, and how can they guide future work on point cloud applications?" *Creative Perspective:* "Consider how these insights could pivot into adjacent fields such as medical imaging or urban mapping." [end sys2] [sys1] To effectively adapt to this problem, a method that uses rotation-invariant features is necessary. This approach can help avoid large-scale data augmentation and maximize efficiency in practical contexts. Focus on integrating spherical voxel methods and leverage adaptive sampling techniques for improved feature extraction. [end sys1]