Huawei Technologies (Sweden)
companyStockholm, Sweden
Research output, citation impact, and the most-cited recent papers from Huawei Technologies (Sweden) (Sweden). Aggregated across the NobleBlocks index of 300M+ scholarly works.
Top-cited papers from Huawei Technologies (Sweden)
Future wireless networks are expected to evolve toward an intelligent and software reconfigurable paradigm enabling ubiquitous communications between humans and mobile devices. They will also be capable of sensing, controlling, and optimizing the wireless environment to fulfill the visions of low-power, high-throughput, massively- connected, and low-latency communications. A key conceptual enabler that is recently gaining increasing popularity is the HMIMOS that refers to a low-cost transformative wireless planar structure comprised of sub-wavelength metallic or dielectric scattering particles, which is capable of shaping electromagnetic waves according to desired objectives. In this article, we provide an overview of HMIMOS communications including the available hardware architectures for reconfiguring such surfaces, and highlight the opportunities and key challenges in designing HMIMOS-enabled wireless communications.
This work presents an estimation of the global electricity usage that can be ascribed to Communication Technology (CT) between 2010 and 2030. The scope is three scenarios for use and production of consumer devices, communication networks and data centers. Three different scenarios, best, expected, and worst, are set up, which include annual numbers of sold devices, data traffic and electricity intensities/efficiencies. The most significant trend, regardless of scenario, is that the proportion of use-stage electricity by consumer devices will decrease and will be transferred to the networks and data centers. Still, it seems like wireless access networks will not be the main driver for electricity use. The analysis shows that for the worst-case scenario, CT could use as much as 51% of global electricity in 2030. This will happen if not enough improvement in electricity efficiency of wireless access networks and fixed access networks/data centers is possible. However, until 2030, globally-generated renewable electricity is likely to exceed the electricity demand of all networks and data centers. Nevertheless, the present investigation suggests, for the worst-case scenario, that CT electricity usage could contribute up to 23% of the globally released greenhouse gas emissions in 2030.
Semantic matching is of central importance to many natural language tasks \cite{bordes2014semantic,RetrievalQA}. A successful matching algorithm needs to adequately model the internal structures of language objects and the interaction between them. As a step toward this goal, we propose convolutional neural network models for matching two sentences, by adapting the convolutional strategy in vision and speech. The proposed models not only nicely represent the hierarchical structures of sentences with their layer-by-layer composition and pooling, but also capture the rich matching patterns at different levels. Our models are rather generic, requiring no prior knowledge on language, and can hence be applied to matching tasks of different nature and in different languages. The empirical study on a variety of matching tasks demonstrates the efficacy of the proposed model on a variety of matching tasks and its superiority to competitor models.
Learning sophisticated feature interactions behind user behaviors is critical in maximizing CTR for recommender systems. Despite great progress, existing methods seem to have a strong bias towards low- or high-order interactions, or require expertise feature engineering. In this paper, we show that it is possible to derive an end-to-end learning model that emphasizes both low- and high-order feature interactions. The proposed model, DeepFM, combines the power of factorization machines for recommendation and deep learning for feature learning in a new neural network architecture. Compared to the latest Wide \& Deep model from Google, DeepFM has a shared input to its "wide" and "deep" parts, with no need of feature engineering besides raw features. Comprehensive experiments are conducted to demonstrate the effectiveness and efficiency of DeepFM over the existing models for CTR prediction, on both benchmark data and commercial data.
The delivery and display of 360-degree videos on Head-Mounted Displays (HMDs) presents many technical challenges. 360-degree videos are ultra high resolution spherical videos, which contain an omnidirectional view of the scene. However only a portion of this scene is displayed on the HMD. Moreover, HMD need to respond in 10 ms to head movements, which prevents the server to send only the displayed video part based on client feedback. To reduce the bandwidth waste, while still providing an immersive experience, a viewport-adaptive 360-degree video streaming system is proposed. The server prepares multiple video representations, which differ not only by their bit-rate, but also by the qualities of different scene regions. The client chooses a representation for the next segment such that its bit-rate fits the available throughput and a full quality region matches its viewing. We investigate the impact of various spherical-to-plane projections and quality arrangements on the video quality displayed to the user, showing that the cube map layout offers the best quality for the given bit-rate budget. An evaluation with a dataset of users navigating 360-degree videos demonstrates that segments need to be short enough to enable frequent view switches.
We address an important problem in sequence-to-sequence (Seq2Seq) learning referred to as copying, in which certain segments in the input sequence are selectively replicated in the output sequence. A similar phenomenon is observable in human language communication. For example, humans tend to repeat entity names or even long phrases in conversation. The challenge with regard to copying in Seq2Seq is that new machinery is needed to decide when to perform the operation. In this paper, we incorporate copying into neural network-based Seq2Seq learning and propose a new model called CopyNet with encoder-decoder structure. CopyNet can nicely integrate the regular way of word generation in the decoder with the new copying mechanism which can choose sub-sequences in the input sequence and put them at proper places in the output sequence. Our empirical study on both synthetic data sets and real world data sets demonstrates the efficacy of CopyNet. For example, CopyNet can outperform regular RNN-based model with remarkable margins on text summarization tasks.
To satisfy the high data demands in future cellular networks, an ultra-densification approach is introduced to shrink the coverage of base station (BS) and improve the frequency reuse.The gain in capacity is expected but at the expense of increased interference, frequent handovers (HOs), increased HO failure (HOF) rates, increased HO delays, increase in ping pong rate, high energy consumption, increased overheads due to frequent HO, high packet losses and bad user experience mostly in high-speed user equipment (UE) scenarios.This paper presents the general concepts of radio access mobility in cellular networks with possible challenges and current research focus.In this article, we provide an overview of HO management in longterm evolution (LTE) and 5G new radio (NR) to highlight the main differences in basic HO scenarios.A detailed literature survey on radio access mobility in LTE, heterogeneous networks (HetNets) and NR is provided.In addition, this paper suggests HO management challenges and enhancing techniques with a discussion on the key points that need to be considered in formulating an efficient HO scheme.
We present sets of spreading sequences that are specifically designed to suit a belief-propagation multiuser detection structure, recently presented for overloaded system scenarios. On one hand, our sequences are of the low-density type; on the other their distance spectrum properties ensure good performance in AWGN channels. Simulations results for raw and coded bit-error probability indicate that significant performance gain is achieved over random sequences and that the loss compared to the single-user bound can be kept small.
In this paper, we propose to learn a deep fitting degree scoring network for monocular 3D object detection, which aims to score fitting degree between proposals and object conclusively. Different from most existing monocular frameworks which use tight constraint to get 3D location, our approach achieves high-precision localization through measuring the visual fitting degree between the projected 3D proposals and the object. We first regress the dimension and orientation of the object using an anchor-based method so that a suitable 3D proposal can be constructed. We propose FQNet, which can infer the 3D IoU between the 3D proposals and the object solely based on 2D cues. Therefore, during the detection process, we sample a large number of candidates in the 3D space and project these 3D bounding boxes on 2D image individually. The best candidate can be picked out by simply exploring the spatial overlap between proposals and the object, in the form of the output 3D IoU score of FQNet. Experiments on the KITTI dataset demonstrate the effectiveness of our framework.
Human computer conversation is regarded as one of the most difficult problems in artificial intelligence. In this paper, we address one of its key sub-problems, referred to as short text conversation, in which given a message from human, the computer returns a reasonable response to the message. We leverage the vast amount of short conversation data available on social media to study the issue. We propose formalizing short text conversation as a search problem at the first step, and employing state-of-the-art information retrieval (IR) techniques to carry out the task. We investigate the significance as well as the limitation of the IR approach. Our experiments demonstrate that the retrieval-based model can make the system behave rather "intelligently", when combined with a huge repository of conversation data from social media.
In this letter we report a new OFDM signalling format characterized by a precoder that renders the emitted signal's phase and amplitude continuous. It achieves superior out-of- band power characteristics at the price of a slightly reduced receiver sensitivity.
In the fifth generation (5G) of mobile broadband systems, radio resource management (RRM) will reach unprecedented levels of complexity. To cope with the ever more sophisticated RRM functionalities and the growing variety of scenarios, while carrying out the prompt decisions required in 5G, this manuscript presents a lean RRM architecture that capitalizes on recent advances in the field of machine learning in combination with the large amount of data readily available in the network from measurements and system observations. The architecture consists of a learner (or a few), which learns RRM policies directly from the data gathered in the network using a single general-purpose learning framework, and a set of distributed actors, which execute RRM policies issued by the learner and repeatedly generate samples of experience. Thus, the complexity of RRM is shifted to the design of the learning framework, while the RRM algorithms derived from this framework are executed in a computationally efficient distributed manner at the radio access nodes. The potential of this approach is verified in a pair of pertinent scenarios, and future directions on applications of machine learning to RRM are discussed. Although we focus on a mobile broadband context, the concepts proposed hereafter extend to any radio access network technology where one can conceive the idea of a central learning unit gathering data from distributed actors.
Millimeter-wave (mmWave) and sub-Terahertz (THz) communications are compelling as an enabler for next-generation wireless networks. In this paper, we study mmWave and sub-THz systems with array-of-subarray architecture. To accommodate the ultrabroad bandwidth in the mmWave and sub-THz bands, time-delay phase shifters are introduced in system design. Our goal is to investigate beamforming training with hybrid processing to extract the dominant channel information, which would fully exploit channel characteristics while respecting the nature of circuit hardware. In particular, codebooks based on time-delay phase shifters are defined and structured. Then, two multi-resolution time-delay codebooks are designed through subarray coordination. One is built on adaptation of physical beam directions, and the other relies on dynamic approximation of beam patterns. Also, a low-complexity system implementation with modifications on the time-delay codebooks is studied. Furthermore, based on the proposed codebooks, a hierarchical beamforming training strategy with reduced overhead is developed to enable simultaneous training for multiple users. Simulation results show that the proposed multi-resolution time-delay codebooks could provide sufficient beam gains and are robust over large bandwidth. Also, the effectiveness of the hierarchical beamforming training is verified.
In this paper, we study the availability of TV white spaces in Europe. Specifically, we focus on the 470-790 MHz UHF band, which will predominantly remain in use for TV broadcasting after the analog-to-digital switch-over and the assignment of the 800 MHz band to licensed services have been completed. The expected number of unused, available TV channels in any location of the 11 countries we studied is 56 percent when we adopt the statistical channel model of the ITU-R. Similarly, a person residing in these countries can expect to enjoy 49 percent unused TV channels. If, in addition, restrictions apply to the use of adjacent TV channels, these numbers reduce to 25 and 18 percent, respectively. These figures are significantly smaller than those recently reported for the United States. We also study how these results change when we use the Longley-Rice irregular terrain model instead. We show that while the overall expected availability of white spaces is essentially the same, the local variability of the available spectrum shows significant changes. This underlines the importance of using appropriate system models before making far-reaching conclusions.
A wideband 2 ×2-slot element for a 60-GHz antenna array is designed by making use of two double-sided printed circuit boards (PCBs). The upper PCB contains the four radiating cavity-backed slots, where the cavity is formed in substrate-integrated waveguide (SIW) using metalized via holes. The SIW cavity is excited by a coupling slot. The excitation slot is fed by a microstrip-ridge gap waveguide formed in the air gap between the upper and lower PCBs. The lower PCB contains the microstrip line, being short-circuited to the ground plane of the lower PCB with via holes, and with additional metalized via holes alongside the microstrip line to form a stopband for parallel-plate modes in the air gap. The designed element can be used in large arrays with distribution networks realized in such microstrip-ridge gap waveguide technology. Therefore, the present paper describes a generic study in an infinite array environment, and performance is measured in terms of the active reflection coefficient S11 and the power lost in grating lobes. The study shows that the radiation characteristics of the array antenna is considerably improved by using a soft surface EBG-type SIW corrugation between each 2 ×2-slot element in E-plane to reduce the mutual coupling. The study is verified by measurements on a 4 ×4 element array surrounded by dummy elements and including a transition to rectangular waveguide WR15.
As the rollout of 4G mobile communication networks takes place, representatives of industry and academia have started to look into the technological developments toward the next generation (5G). Several research projects involving key international mobile network operators, infrastructure manufacturers, and academic institutions, have been launched recently to set the technological foundations of 5G. However, the architecture of future 5G systems, their performance, and mobile services to be provided have not been clearly defined. In this paper, we put forth the vision for 5G as the convergence of evolved versions of current cellular networks with other complementary radio access technologies. Therefore, 5G may not be a single radio access interface but rather a "network of networks". Evidently, the seamless integration of a variety of air interfaces, protocols, and frequency bands, requires paradigm shifts in the way networks cooperate and complement each other to deliver data rates of several Gigabits per second with end-to-end latency of a few milliseconds. We provide an overview of the key radio technologies that will play a key role in the realization of this vision for the next generation of mobile communication networks. We also introduce some of the research challenges that need to be addressed.
Accuracy of depth estimation from static images has been significantly improved recently, by exploiting hierarchical features from deep convolutional neural networks (CNNs). Compared with static images, vast information exists among video frames and can be exploited to improve the depth estimation performance. In this work, we focus on exploring temporal information from monocular videos for depth estimation. Specifically, we take the advantage of convolutional long short-term memory (CLSTM) and propose a novel spatial-temporal CSLTM (ST-CLSTM) structure. Our ST-CLSTM structure can capture not only the spatial features but also the temporal correlations/consistency among consecutive video frames with negligible increase in computational cost. Additionally, in order to maintain the temporal consistency among the estimated depth frames, we apply the generative adversarial learning scheme and design a temporal consistency loss. The temporal consistency loss is combined with the spatial loss to update the model in an end-to-end fashion. By taking advantage of the temporal information, we build a video depth estimation framework that runs in real-time and generates visually pleasant results. Moreover, our approach is flexible and can be generalized to most existing depth estimation frameworks. Code is available at: https://tinyurl.com/STCLSTM
This paper is concerned with the channel estimation problem in millimeter wave (mmWave) wireless systems with large antenna arrays. By exploiting the inherent sparse nature of the mmWave channel, we first propose a fast channel estimation (FCE) algorithm based on a novel overlapped beam pattern design, which can increase the amount of information carried by each channel measurement compared to the existing nonoverlapped designs and thus reduce the required channel estimation time. We develop a maximum likelihood estimator to optimally extract the path information from the channel measurements. Then, we propose a novel rate-adaptive channel estimation (RACE) algorithm, which can dynamically adjust the number of channel measurements based on the expected probability of estimation error (PEE). The performance of both proposed algorithms is analyzed. For the FCE algorithm, an approximate closed-form expression for the PEE is derived. For the RACE algorithm, a lower bound for the minimum signal energy-to-noise ratio required for a given number of channel measurements is developed based on the Shannon-Hartley theorem. Simulation results show that the FCE algorithm significantly reduces the number of channel estimation measurements compared to the existing algorithms using nonoverlapped beam patterns. By adopting the RACE algorithm, we can achieve up to a 6 dB gain in signal energy-to-noise ratio for the same PEE compared to the existing algorithms.
Age estimation is an important yet very challenging problem in computer vision. Existing methods for age estimation usually apply a divide-and-conquer strategy to deal with heterogeneous data caused by the non-stationary aging process. However, the facial aging process is also a continuous process, and the continuity relationship between different components has not been effectively exploited. In this paper, we propose BridgeNet for age estimation, which aims to mine the continuous relation between age labels effectively. The proposed BridgeNet consists of local regressors and gating networks. Local regressors partition the data space into multiple overlapping subspaces to tackle heterogeneous data and gating networks learn continuity aware weights for the results of local regressors by employing the proposed bridge-tree structure, which introduces bridge connections into tree models to enforce the similarity between neighbor nodes. Moreover, these two components of BridgeNet can be jointly learned in an end-to-end way. We show experimental results on the MORPH II, FG-NET and Chalearn LAP 2015 datasets and find that BridgeNet outperforms the state-of-the-art methods.
Massive multiple-input, multiple-output (MIMO) is seen as an enabling technology to fulfill dramatic improvements in spectral efficiency for fifth-generation (5G) deployment in 2020. For massive MIMO systems, the learning loop from early-stage prototype design to final-stage performance validation is expected to be slow and ineffective. There is a strong need to evaluate massive MIMO base station (BS) performance with over-the-air (OTA) methods. Until now, such OTA solutions have not been discussed for massive MIMO BS systems.