Sign In to Follow Application
View All Documents & Correspondence

Cluster System, Cluster System Control Method, Server Device, Control Method, And Non Transitory Computer Readable Medium Having Program Stored Therein

Abstract: The present invention reliably determines as to whether or not a service is properly provided to a client by an active-system server. A cluster system (1) is provided with: an active-system server (2) for providing a prescribed service to a client device via a network (4); and a standby-system server (3) which, in the event that a fault occurs in the active-system server (2), provides the prescribed service to the client device in place of the active-system server (2). The standby-system server (3) has a monitoring unit (6) that accesses, via a network (4), the prescribed service provided by the active-system server (2), and monitors whether or not the service can be accessed correctly. The active-system server (2) has a cluster control unit (5) that implements failover in the case when the monitoring unit (6) of the standby-system server (3) has determined that the prescribed service provided by the active-system server (2) cannot be accessed correctly.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
13 March 2020
Publication Number
35/2020
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
archana@anandandanand.com
Parent Application

Applicants

NEC CORPORATION
7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Inventors

1. OSAWA Ryosuke
c/o NEC Corporation, 7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Specification

Specification
Title of invention: Cluster system, cluster system control method, server device, control method, and non-transitory computer-readable medium storing a program
Technical field
[0001]
 The present invention relates to a cluster system, a cluster system control method, a server device, a control method, and a non-transitory computer-readable medium in which a program is stored.
Background technology
[0002]
 A HA (High Availability) cluster system is a technique for improving system availability. For example, Japanese Patent Laid-Open No. 2004-242242 discloses an active system device that executes a process job in response to a process request from a client, a standby system device that takes over the process job when the active system device fails, an active system device and a standby system. Disclosed is a cluster system having a LAN (Local Area Network) that connects a device and a client, and a communication path that connects between an active device and a standby device.
[0003]
 In the HA cluster system, generally, there are an active server that provides a predetermined service such as a business service and a standby server that takes over the service and provides the service when a failure occurs. Each of the running servers forming the cluster mutually monitors whether or not they can communicate with each other. That is, monitoring by heartbeat. In addition to this, the active server monitors whether the local server can provide services normally, and the standby server also monitors whether the local server can normally take over services. doing.
Prior art documents
Patent literature
[0004]
Patent Document 1: Japanese Patent Laid-Open No. 11-338725
Summary of the invention
Problems to be Solved by the Invention
[0005]
 The active server performs, for example, disk monitoring, NIC (Network Interface Card) monitoring, public LAN monitoring, and monitoring for specific services (HTTP (Hypertext Transfer Protocol) protocol monitoring, etc.) in monitoring of its own server. By combining them, it is judged whether the own server can properly provide the service. However, since this determination is made by the active server itself, it cannot be assured that the service is actually provided to the external client via the public LAN.
[0006]
 In view of the above-described problems, an object of the present invention is to provide a cluster system, a cluster system control method, and a server that can more reliably determine whether or not a service is properly provided to a client by an active server. An object is to provide a non-transitory computer-readable medium in which a device, a control method, and a program are stored.
Means for solving the problem
[0007]
 A cluster system according to an aspect of the present invention replaces an active server device that provides a predetermined service to a client device via a network and the active server device when an abnormality occurs in the active server device. A standby system server device that provides the predetermined service to the client device, wherein the standby system server device accesses the predetermined service provided by the active system server device via the network, It has a first monitoring means for monitoring whether or not it can be normally accessed, and if the active server device cannot normally access the predetermined service provided by the active server device, It has a cluster control means for carrying out a failover when it is judged by the first monitoring means.
[0008]
 In a control method of a cluster system according to one aspect of the present invention, an active server device provides a predetermined service to a client device via a network, and a standby server device that constitutes a cluster system together with the active server, Through the network, the predetermined service provided by the active server device is accessed and monitored for normal access, and the active server device provides the predetermined service provided by the active server device. If the standby server device determines that the service cannot be normally accessed, failover is performed.
[0009]
 A server device according to an aspect of the present invention normally accesses a service providing unit that provides a predetermined service to a client device via a network and normally accesses the predetermined service provided by the service providing unit via the network. If the monitoring result sent by the standby server device that monitors whether or not it can be accessed is acquired, and if the monitoring result indicates that the standby server device cannot normally access the predetermined service, failover is performed. The standby server device is a device that takes over the provision of the predetermined service to the client device when a failover is performed.
 Further, in the control method according to an aspect of the present invention, a standby is provided in which a predetermined service is provided to a client device via a network, and the predetermined service is accessed via the network and monitored for normal access. When the monitoring result transmitted by the system server device is acquired and the monitoring result indicates that the standby server device cannot normally access the predetermined service, failover is performed, and the standby server device It is a device that takes over the provision of the predetermined service to the client device when a failover is performed.
[0010]
 A program according to an aspect of the present invention accesses a service providing step of providing a predetermined service to a client device via a network, and accesses the predetermined service provided by the processing of the service providing step via the network. If the monitoring result transmitted by the standby server device that monitors whether or not the standby server can be normally accessed is acquired, and the monitoring result indicates that the standby server device cannot normally access the predetermined service, a failover occurs. And a standby control server device that takes over the provision of the predetermined service to the client device when a failover is performed.
Effect of the invention
[0011]
 Advantageous Effects of Invention According to the present invention, a cluster system, a cluster system control method, a server device, a control method, and a cluster system capable of more reliably determining whether or not a service is appropriately provided to a client by an active server, A non-transitory computer-readable medium in which a program is stored can be provided.
Brief description of the drawings
[0012]
FIG. 1 is a block diagram showing an example of a configuration of a cluster system according to the outline of the embodiment.
FIG. 2 is a block diagram showing an example of a functional configuration of the cluster system according to the embodiment.
FIG. 3 is a block diagram showing an example of a hardware configuration of each server that constitutes the cluster system according to the exemplary embodiment.
FIG. 4 is a sequence chart showing an operation example at the start of providing a business service in a cluster system.
FIG. 5 is a sequence chart showing an operation example of the cluster system when an abnormality of a business service is detected in a standby server.
FIG. 6 is a sequence chart showing an operation example when an abnormality occurs in one standby server in the cluster system.
FIG. 7 is a sequence chart showing an operation example when an error occurs in all standby servers in the cluster system.
FIG. 8 is a block diagram showing an example of a configuration of a server device according to the embodiment.
MODE FOR CARRYING OUT THE INVENTION
[0013]
 For clarity of explanation, the following description and drawings are appropriately omitted and simplified. In each drawing, the same elements are denoted by the same reference numerals, and redundant description is omitted as necessary.
[0014]

 Prior to description of the embodiment, an outline of the embodiment according to the present invention will be described. FIG. 1 is a block diagram showing an example of the configuration of a cluster system 1 according to the outline of the embodiment. As shown in FIG. 1, the cluster system 1 has an active server 2, a standby server 3, and a network 4.
[0015]
 The active server 2 is a server device that provides a predetermined service to a client device (not shown) via the network 4. That is, the client device accesses a predetermined service provided by the active server 2 via the network 4.
 The standby server 3 is a server device that provides a predetermined service to the client device in place of the active server 2 when an abnormality occurs in the active server 2.
[0016]
 The standby server 3 has a monitoring unit 6 (monitoring unit), and the active server 2 has a cluster control unit 5 (cluster control unit). The monitoring unit 6 accesses a predetermined service provided by the active server 2 via the network 4 and monitors whether or not the service can be normally accessed. That is, the monitoring unit 6 accesses the active server 2 via the network 4, like the client device. When the monitoring unit 6 of the standby server 3 determines that the predetermined service provided by the active server 2 cannot be normally accessed, the cluster control unit 5 performs a failover. The cluster control unit 5 executes, for example, failover processing so that the standby server 3 can take over the provision of a predetermined service.
[0017]
 Generally, the loopback address is used when the active server itself monitors the service. Therefore, the monitoring is performed by the communication processing closed in the server. Therefore, it is not possible to confirm whether or not the service can be accessed by actually communicating with the specific port number via the network used by the client device. In addition, although it is possible to confirm the communication by ping (ICMP (Internet Control Message Protocol)) to the network device connected to the server, it is specified by the failure of the external network device, the bug of OS (Operating System), the setting mistake of the firewall, etc. If external communication is not possible with the port number of, it is difficult to detect the abnormality. For this reason, it cannot be reliably determined that the service can be provided to the client device. Further, it is possible to expect a more reliable judgment by introducing the operation management software and the operation management server for monitoring the service, but the introduction and operation costs of these are required.
[0018]
 On the other hand, in the cluster system 1, the standby server 3 accesses the service provided by the active server 2 via the same network 4 that the client device uses for access, and thus the active server 2 provides the service. Monitor. Therefore, the service provision can be monitored by the same access as the client device that actually receives the service provision.
[0019]
 Therefore, according to the cluster system 1, it is possible to more reliably determine whether or not the service is appropriately provided to the client device by the active server. Further, since the standby system server 3 is used for monitoring, it is not necessary to newly prepare an operation management server for monitoring the service or newly install operation management software for monitoring the service, and the cost of installation and operation is reduced. Can be suppressed.
[0020]

 Hereinafter, an embodiment of the present invention will be described. FIG. 2 is a block diagram showing an example of the functional configuration of the cluster system 10 according to the embodiment. FIG. 3 is a block diagram showing an example of the hardware configuration of each server that constitutes the cluster system 10.
[0021]
 As shown in FIG. 2, the cluster system 10 according to the present embodiment includes an active server 100A, a standby server 100B, a standby server 100C, a network 200, and a network 300. The active server 100A and the standby servers 100B and 100C have clusterware 110A, 110B and 110C, respectively, and by communicating with each other via the networks 200 and 300, an HA cluster system is configured. In the following description, the servers that make up the cluster system 10 may be referred to as the servers 100 when they are referred to without particular distinction.
[0022]
 The active server 100A corresponds to the active server 2 of FIG. 1, and is a server that provides business services to clients via the network 200. In addition, the standby servers 100B and 100C correspond to the standby server 3 of FIG. 1, and when an abnormality occurs in the active server 100A, provide the business service to the client in place of the active server 100A. It is a server. That is, the standby servers 100B and 100C are devices that take over the provision of the business service to the client when a failover is performed.
[0023]
 As shown in FIG. 2, the active server 100A has a business service providing unit 120A and clusterware 110A. Further, the clusterware 110A includes a business service control unit 111A, another server monitoring unit 112A, an own server monitoring unit 113A, and a cluster control unit 114A. The standby servers 100B and 100C have the same configuration as the active server 100A. That is, the standby server 100B has a business service providing unit 120B, a business service control unit 111B, another server monitoring unit 112B, a self server monitoring unit 113B, and clusterware 110B including a cluster control unit 114B. Further, the standby server 100C includes a business service providing unit 120C, a business service control unit 111C, another server monitoring unit 112C, an own server monitoring unit 113C, and clusterware 110C including a cluster control unit 114C.
[0024]
 Note that the business service providing units 120A, 120B, and 120C may be referred to as the business service providing unit 120 if they are referred to without distinction. The clusterware 110A, 110B, and 110C may be referred to as clusterware 110 when they are referred to without distinction. When referring to the business service control units 111A, 111B, and 111C without making a distinction, they may be referred to as the business service control unit 111. When referring to other server monitoring units 112A, 112B, and 112C without making a distinction, they may be referred to as other server monitoring unit 112. When referring to the own server monitoring units 113A, 113B, and 113C without making a distinction, they may be referred to as the own server monitoring unit 113. The cluster control units 114A, 114B, 114C may be referred to as the cluster control unit 114 when they are referred to without particular distinction.
[0025]
 Here, an example of the hardware configuration of each server 100 is shown with reference to FIG. The server 100 includes a network interface 151, a memory 152, and a processor 153.
[0026]
 The network interface 151 is used to communicate with other devices via the network 200 or the network 300. The network interface 151 may include, for example, a network interface card (NIC).
[0027]
 The memory 152 is composed of a combination of a volatile memory and a non-volatile memory. Memory 152 may include storage located remotely from processor 153. In this case, the processor 153 may access the memory 152 via an input/output interface (not shown).
[0028]
 The memory 152 is used to store software (computer program) executed by the processor 153 and the like.
[0029]
 This program can be stored using various types of non-transitory computer readable media, and can be supplied to a computer. Non-transitory computer readable media include various types of tangible storage media. Examples of non-transitory computer readable media are magnetic recording media (eg flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (eg magneto-optical disks), Compact Disc Read Only Memory (CD-ROM), CD- R, CD-R/W, semiconductor memory (for example, mask ROM, Programmable ROM (PROM), Erasable PROM (EPROM), flash ROM, Random Access Memory (RAM)) are included. In addition, the program may be supplied to the computer by various types of transitory computer readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable medium can supply the program to the computer via a wired communication path such as an electric wire and an optical fiber, or a wireless communication path.
[0030]
 The processor 153 reads the computer program from the memory 152 and executes the computer program to perform the process of the business service providing unit 120, the process of the clusterware 110, and other processes. The processor 153 may be, for example, a microprocessor, MPU, or CPU. The processor 153 may include multiple processors.
[0031]
 The network 200 is a public LAN and is used for mutual communication between the servers 100 and communication with external clients. That is, the network 200 is used as a network path for providing business services to external clients.
[0032]
 The network 300 is an interconnect LAN and is used for mutual communication between the servers 100, but is not used for communication with external clients. The network 300 is used as a dedicated line in the cluster system 10 in consideration of avoiding influence on business services and security. The network 300 is used for internal communication in the cluster system 10 (processing request, heartbeat between each server 100 (life monitoring), synchronization of business data, etc.).
[0033]
 As described above, the network 200 is a network different from the network 300 used for performing the life and death monitoring between the active server 100 and the standby server 100.
[0034]
 Next, the configuration of each server 100 shown in FIG. 2 will be described.
 The business service providing unit 120 receives an access via the network 200 and provides a predetermined business service. That is, the business service providing unit 120 (service providing unit) provides a predetermined service to the client device via the network 200. The business service providing unit 120 is a module that operates in the active server 100. Therefore, the business service control unit 111A in the active server 100A is operating, but the business service control units 111B and 111C in the standby servers 100B and 100C are not operating.
[0035]
 The cluster control unit 114 corresponds to the cluster control unit 5 in FIG. 1, cooperates with the cluster control unit 114 of another server, and performs various controls for realizing the availability of the cluster system 10. The cluster control unit 114 performs alive monitoring of other servers 100 based on a heartbeat, execution of failover, and the like. Further, the cluster control unit 114 notifies the other server 100 of the monitoring result of the other server monitoring unit 112 and realizes the synchronization of the monitoring result among the servers 100. The monitoring result of the business service synchronized with each server 100 is used as a display of a cluster management GUI (Graphical User Interface) of each server 100 for determining whether the business service is normal. The other processing contents of the cluster control unit 114 will be described later together with the operation of the cluster system 10.
[0036]
 The business service control unit 111 controls starting and stopping of the business service providing unit 120. In the present embodiment, the business service control unit 111 controls to start the business service providing unit 120 in response to a start request from the cluster control unit 114, and in response to a stop request from the cluster control unit 114, a business service is requested. The service providing unit 120 is controlled to stop. That is, for example, the business service control unit 111A controls the business service providing unit 120A to be activated in response to a start request from the cluster control unit 114A, and provides the business service in response to a stop request from the cluster control unit 114A. The part 120A is controlled to stop.
[0037]
 The own server monitoring unit 113 (monitoring means) monitors the status of the disk, NIC, etc. of the own server. When the self-server monitoring unit 113 detects a failure by monitoring, the self-server monitoring unit 113 notifies the cluster control unit 114 of the abnormality. That is, for example, the own server monitoring unit 113A monitors the operating state of the active server 100A itself and notifies the cluster control unit 114A of the monitoring result. Similarly, for example, the own server monitoring unit 113B monitors the operating status of the standby server 100B itself and notifies the cluster control unit 114B of the monitoring result.
[0038]
 The other server monitoring unit 112 (monitoring unit) corresponds to the monitoring unit 6 of FIG. 1 and accesses a predetermined service provided by the business service providing unit 120 of the active server 100 via the network 200 Monitor for normal access. That is, for example, the other server monitoring unit 112B accesses a predetermined service provided by the business service providing unit 120A and monitors whether or not the service can be normally accessed. The other server monitoring unit 112 is a module that operates in the standby server 100. Therefore, the other server monitoring units 112B and 112C of the standby servers 100B and 100C are operating, but the other server monitoring unit 112 of the active server 100A is not operating. That is, the standby servers 100B and 100C monitor whether the business services provided by the active server 100A can be accessed via the network 200 by the other server monitoring units 112B and 112C. In addition, in this Embodiment, the other server monitoring part 112 monitors regularly. The other server monitoring unit 112 notifies the cluster control unit 114 of the monitoring result. That is, for example, the other server monitoring unit 112B notifies the cluster control unit 114B of the monitoring result.
[0039]
 The other server monitoring unit 112 performs a monitoring process according to the business service protocol (FTP, HTTP, IMAP4, POP3, SMTP, etc.) provided by the active server 100. Further, the other server monitoring unit 112 performs the monitoring process via the network 200 as described above in order to perform the same access as the actual external client. Here, a specific example of the monitoring process by the other server monitoring unit 112 will be described.
[0040]
 When the business service provided is a service using FTP (File Transfer Protocol), that is, when the active server 100A functions as an FTP server, the other server monitoring unit 112 connects to the monitored FTP server and the user Perform authentication process. After that, the other server monitoring unit 112 acquires the file list of the FTP server. The other server monitoring unit 112 determines that the service provision is normally performed because all of these processes are normal.
[0041]
 When the business service provided is a service using HTTP, that is, when the active server 100A functions as an HTTP server, the other server monitoring unit 112 transmits an HTTP request to the HTTP server to be monitored, and the HTTP Since the processing result of the HTTP response from the server is normal, it is determined that the service is normally provided.
[0042]
 When the business service provided is a service using IMAP4 (Internet Message Access Protocol 4), that is, when the active server 100A functions as an IMAP server, the other server monitoring unit 112 connects to the IMAP server to be monitored. , Perform user authentication process. After that, the other server monitoring unit 112 executes the NOOP command. The other server monitoring unit 112 determines that the service provision is normally performed because all of these processes are normal.
[0043]
 When the business service provided is a service using POP3 (Post Office Protocol 3), that is, when the active server 100A functions as a POP3 server, the other server monitoring unit 112 connects to the POP3 server to be monitored, Perform user authentication process. After that, the other server monitoring unit 112 executes the NOOP command. The other server monitoring unit 112 determines that the service provision is normally performed because all of these processes are normal.
[0044]
 When the business service provided is a service using SMTP (Simple Mail Transfer Protocol), that is, when the active server 100A functions as an SMTP server, the other server monitoring unit 112 connects to the SMTP server to be monitored, Perform user authentication process. After that, the other server monitoring unit 112 executes the NOOP command. The other server monitoring unit 112 determines that the service provision is normally performed because all of these processes are normal.
[0045]
 Note that a timeout time or the number of retries may be provided as a threshold for determining an abnormality in the monitoring by the other server monitoring unit 112 so that optimal monitoring can be realized according to the system environment. For example, the other server monitoring unit 112 may determine that the service is not normally provided by the active server 100 when the business service cannot be normally accessed within a predetermined timeout time. In addition, the other server monitoring unit 112 may determine that the service is not normally provided by the active server 100 when the business service cannot be normally accessed within a predetermined number of retries.
[0046]
 With the configuration as described above, the cluster system 10 performs, for example, the following operation. In the active server 100A, the business service providing unit 120A is activated under the control of the business service control unit 111A at the request of the cluster control unit 114A. As a result, the business service providing unit 120A provides business services to external clients via the network 200. Further, in the active server 100A, the own server monitoring unit 113 monitors the operating status of the active server 100A itself, and when a failure occurs, notifies the cluster control unit 114A of the abnormality. The cluster control unit 114A that has received the abnormality notification requests the business service control unit 111A to stop the business service, thereby stopping the operation of the business service providing unit 120A. After that, the cluster control unit 114A requests the cluster control unit 114B of the standby server 100B to start the business service providing unit 120B, and the business server can be provided from the standby server 100B to perform failover.
[0047]
 In the standby system servers 100B and 100C, the own server monitoring units 113B and 113C monitor the operating status of the own server. Further, the other server monitoring units 112B and 112C monitor whether or not the business service provided by the active server 100A can be accessed via the network 200. The cluster control units 114B and 114C notify the result of monitoring by the other server monitoring unit 112 to the cluster control unit 114A of the active server 100A. When the cluster control unit 114A obtains a monitoring result that the business service of the active server 100A cannot be accessed in both the standby servers 100B and 100C, it determines that a failure has occurred in the active server 100A, Perform business service failover.
[0048]
 As described above, in the present embodiment, the cluster control unit 114A is determined by the other server monitoring unit 112 of the plurality of standby servers 100 to be unable to normally access the predetermined service provided by the active server 100A. If so, perform a failover. More specifically, the cluster control unit 114A has a predetermined service provided by the active server 100A by the other server monitoring units 112 of the standby servers 100 of a predetermined ratio or more among the plurality of standby servers 100. If it is determined that the normal access cannot be made, failover is performed. In the present embodiment, specifically, when it is determined that the service cannot be normally accessed by the other server monitoring unit 112 of the majority of standby servers 100 of the plurality of standby servers 100, The cluster control unit 114A of the active server 100A executes failover. As described above, in the present embodiment, the cluster system 10 integrates the monitoring results of the other server monitoring units 112 of the plurality of standby servers 100 to determine whether or not to perform failover. Therefore, the occurrence of failover due to the false detection of the other server monitoring unit 112 due to the failure of the standby server 100 is suppressed.
[0049]
 Next, a specific operation example of the cluster system 10 will be described using a sequence chart. FIG. 4 is a sequence chart showing an operation example at the start of providing the business service in the cluster system 10. The operation of the cluster system 10 will be described below with reference to FIG.
[0050]
 In step 101 (S101), the cluster control unit 114A requests the business service control unit 111A to activate the business service providing unit 120A. Therefore, in step 102 (S102), the business service control unit 111A activates the business service providing unit 120A.
[0051]
 When the business service becomes available, in step 103 (S103), the cluster control unit 114A requests the cluster control unit 114B to start the regular monitoring process for the business service that is started to be provided by the active server 100A. .. For this reason, in step 104 (S104), the cluster control unit 114B requests the other server monitoring unit 112B to start the regular monitoring process of the business service started to be provided by the active server 100A.
[0052]
 Next, in step 105 (S105), the cluster control unit 114A requests the cluster control unit 114C to start the regular monitoring process for the business service that has been provided by the active server 100A. Therefore, in step 106 (S106), the cluster control unit 114C requests the other server monitoring unit 112C to start the regular monitoring process of the business service that is started to be provided by the active server 100A.
[0053]
 Next, in step 107 (S107), the other server monitoring unit 112B executes a regular monitoring process of the business service. The other server monitoring unit 112B actually accesses the business service via the network 200 and confirms whether it is available or not. It is assumed here that it is determined that normal access is possible (that is, the business service is normally available).
 In step 108 (S108), the other server monitoring unit 112B notifies the cluster control unit 114B of the monitoring result (normal) performed in step 107.
 In step 109 (S109), the cluster control unit 114B notifies each of the other servers 100 of the monitoring result (normal) notified in step 108, and synchronizes the monitoring result.
[0054]
 Next, in step 110 (S110), the cluster control unit 114A checks the synchronized monitoring results and determines whether failover is necessary. Here, the cluster control unit 114A determines that failover is unnecessary.
[0055]
 Next, in step 111 (S111), the other server monitoring unit 112C carries out a regular monitoring process of the business service, similarly to the other server monitoring unit 112B. It is assumed here that it is determined that normal access is possible (that is, the business service is normally available).
 In step 112 (S112), the other server monitoring unit 112C notifies the cluster control unit 114C of the monitoring result (normal) performed in step 111.
 In step 113 (S113), the cluster control unit 114C notifies each of the other servers 100 of the monitoring result (normal) notified in step 112, and synchronizes the monitoring result.
[0056]
 Next, in step 114 (S114), the cluster control unit 114A confirms the synchronized monitoring result and determines whether failover is necessary. Here, the cluster control unit 114A determines that failover is unnecessary.
[0057]
 FIG. 5 is a sequence chart showing an operation example of the cluster system 10 when the standby server 100 detects an abnormality in the business service. The operation of the cluster system 10 when the other server monitoring unit 112 detects an abnormality will be described below with reference to FIG.
[0058]
 In step 201 (S201), a failure occurs in the business service provided by the business service providing unit 120A, and the business service becomes unavailable from the external client.
[0059]
 In step 202 (S202), the other server monitoring unit 112B executes the regular monitoring process of the business service, as in step 107 of FIG. In step 202, the other server monitoring unit 112 determines that the normal service cannot be accessed (that is, the business service cannot be normally used).
 In step 203 (S203), the other server monitoring unit 112B notifies the cluster control unit 114B of the monitoring result (abnormality) performed in step 202.
 In step 204 (S204), the cluster control unit 114B notifies each of the other servers 100 of the monitoring result (abnormality) notified in step 203, and synchronizes the monitoring result.
[0060]
 Next, in step 205 (S205), the cluster control unit 114A checks the synchronized monitoring result and determines whether failover is necessary. At present, the number of standby system servers 100 that have detected an abnormality is one, which is less than the majority of the total number of standby system servers 100. Therefore, the cluster control unit 114A determines that failover is unnecessary.
[0061]
 In step 206 (S206), the other server monitoring unit 112C executes the regular monitoring process of the business service, as in step 111 of FIG. In step 206, the other server monitoring unit 112 determines that the normal access cannot be made (that is, the business service cannot be normally used).
 In step 207 (S207), the other server monitoring unit 112C notifies the cluster control unit 114C of the monitoring result (abnormality) performed in step 206.
 In step 208 (S208), the cluster control unit 114C notifies each of the other servers 100 of the monitoring result (abnormality) notified in step 207, and synchronizes the monitoring result.
[0062]
 Next, in step 209 (S209), the cluster control unit 114A confirms the synchronized monitoring result and determines whether failover is necessary. The number of standby system servers 100 that have detected an abnormality is two, which is more than the majority of the total number of standby system servers 100. Therefore, the cluster control unit 114A starts failover. Specifically, the following processing is performed as failover processing.
[0063]
 In step 210 (S210), the cluster control unit 114A requests the cluster control unit 114B to end the regular monitoring process of the business service. Therefore, in step 211 (S211), the cluster control unit 114B requests the other server monitoring unit 112B to end the regular monitoring process of the business service.
[0064]
 In step 212 (S212), the cluster control unit 114A requests the cluster control unit 114C to end the regular monitoring process of the business service. Therefore, in step 213 (S213), the cluster control unit 114C requests the other server monitoring unit 112C to end the regular monitoring process of the business service.
[0065]
 Next, in step 214 (S214), the cluster control unit 114A requests the business service control unit 111A to stop providing the business service. Therefore, in step 215 (S215), the business service control unit 111A stops the processing of the business service providing unit 120A. After that, the startup processing is performed in any of the standby servers 100, and the failover is completed. That is, the business server is handed over to the standby server 100.
[0066]
 The first embodiment has been described above. In the present embodiment, as described above, the cluster control unit 114A of the active server 100A accesses the predetermined service provided by the business service providing unit 120A via the network 200 and monitors whether or not the service can be normally accessed. The monitoring results transmitted by the standby servers 100B and 100C are acquired. Then, when the acquired monitoring result indicates that the standby server 100 cannot normally access the predetermined service, the cluster control unit 114A performs a failover. Therefore, the service provision is monitored by the same access as the client who actually receives the service provision. Therefore, according to the cluster system 10 according to the first embodiment, it is possible to more reliably determine whether or not the service is appropriately provided to the client by the active server 100A. Further, since the standby system servers 100B and 100C are used for monitoring, there is no need to newly prepare an operation management server for monitoring a service or newly install service management operation management software.
[0067]
 Further, in the present embodiment, when it is determined by the other server monitoring unit 112 of a majority of the standby system servers 100 of the plurality of standby system servers 100 that the service cannot be normally accessed, the active server 100A of the active system servers 100A. The cluster control unit 114A executes failover. Therefore, it is possible to suppress the influence of erroneous detection when the other server monitoring unit 112 is not normally monitored due to a failure of the standby server 100 or a failure of a network device connected to the standby server 100. it can.
[0068]
 
 Next, the difference between Embodiment 2 and Embodiment 1 will be described. In the first embodiment, the cluster control unit 114A uses the other server monitoring units 112 of the standby system servers 100 of a predetermined ratio or more among all the standby system servers 100 configuring the cluster system 10 to cause the active system servers to operate. When it is determined that the predetermined service provided by 100A cannot be normally accessed, failover is performed. That is, in the first embodiment, the monitoring result of the other server monitoring unit 112 of the standby server 100 is used for determining whether to perform failover regardless of whether the standby server 100 is operating normally. I was there.
[0069]
 On the other hand, in the present embodiment, the cluster control unit 114A of the active server 100A waits for a predetermined number or more of the standby servers 100 that have not detected an abnormality by the local server monitoring unit 113. When the other server monitoring unit 112 of the system server 100 determines that the predetermined service provided by the active server 100A cannot be normally accessed, failover is performed. That is, in the present embodiment, when the local server monitoring unit 113 of the standby server 100 detects an abnormality of the local server, the monitoring result of the other server monitoring unit 112 of the server 100 is the same as that of the majority decision. Not included in the number of cases.
[0070]
 In addition, in the present embodiment, the cluster control unit 114A of the active server 100A, if there is no standby server 100 in which an error is not detected by the local server monitoring unit 113, the local server monitoring unit of the active server 100A. Based on the monitoring result by 113A, it is determined whether or not to perform failover. That is, in the present embodiment, when there is no standby system server 100 that is determined to be operating normally by the local server monitoring unit 113, in other words, the other server monitoring unit 112 operates normally. When there is no standby system server 100, the cluster control unit 114A of the active server 100A determines whether the business service is normally provided based on the monitoring result of the own server monitoring unit 113A. The own server monitoring unit 113A of the active server 100A accesses the service provided by the business service control unit 111A by using, for example, the loopback address, thereby performing the monitoring process for the business service.
[0071]
 A specific operation example of the cluster system 10 according to the second exemplary embodiment will be described using a sequence chart. FIG. 6 is a sequence chart showing an operation example when an abnormality occurs in one standby server 100 in the cluster system 10 according to the second exemplary embodiment. The operation of the cluster system 10 will be described below with reference to FIG. Note that the example shown in FIG. 6 shows a case where an abnormality has occurred in the standby server 100B of the two standby servers 100 configuring the cluster system 10. The sequence chart shown in FIG. 6 is, for example, a sequence chart following the sequence chart shown in FIG.
[0072]
 In step 301 (S301), a failure occurs in the standby server 100B, and the failure is detected by the monitoring process of the local server monitoring unit 113B of the standby server 100B.
 In step 302 (S302), the own server monitoring unit 113B notifies the cluster control unit 114B of the monitoring result (abnormality) performed in step 301.
 In step 303 (S303), the cluster control unit 114B notifies each of the other servers 100 of the monitoring result (abnormality) notified in step 302, and synchronizes the monitoring result.
[0073]
 Next, in step 304 (S304), the cluster control unit 114A confirms the synchronized monitoring result. Since the synchronized monitoring result indicates abnormality of the standby server 100B, the cluster control unit 114A excludes the monitoring result of the other server monitoring unit 112B of the standby server 100B from the failover determination (exclusion flag). ).
[0074]
 Since setting the exclusion flag of the standby server 100B may change the determination as to whether or not failover should be performed, in step 305 (S305), the cluster control unit 114A causes the standby server 100B The monitoring result of the other server monitoring unit 112 is reconfirmed. At this point, it is assumed that no abnormality has been detected in any of the standby server 100 and the other server monitoring unit 112. In this case, the cluster control unit 114A determines that failover is unnecessary.
[0075]
 On the other hand, in step 306 (S306), the cluster control unit 114B of the standby server 100B suspends the monitoring process of the other server monitoring unit 112B. The monitoring process of the other server monitoring unit 112B that is temporarily stopped here does not resume until the state of the standby server 100B returns to normal.
[0076]
 As described above, the cluster control unit 114A sets the exclusion flag when the standby server 100 notifies that the server 100 is abnormal. In the second embodiment, the cluster control unit 114A uses the monitoring result of the other server monitoring unit 112 of the standby system servers 100 that do not have the exclusion flag set among the standby system servers 100 that configure the cluster system 10. , Determine whether to implement failover. That is, the number of standby system servers 100 that make up the cluster system 10 is N (N is an integer of 1 or more), and the number of standby system servers 100 for which the exclusion flag is not set is n 1 (n 1 is an integer of 1 or more and N or less). Further, among the n 1 standby servers 100, the number of servers in which the other server monitoring unit 112 detects an abnormality is n 2 (n 2 is an integer of 1 or more and n 1 or less). In this case, the cluster control unit 114A performs failover when n 2 /n 1 is equal to or higher than a predetermined ratio (for example, n 2 is a majority of n 1 ).
[0077]
 As described above, in the present embodiment, when the local server monitoring unit 113 of the standby server 100 detects an abnormality of the local server, the monitoring result of the other server monitoring unit 112 of the server 100 is failed over. It is not taken into consideration in the decision of implementation. For this reason, it is possible to suppress the influence on the determination of the execution of failover due to an erroneous monitoring result by the other server monitoring unit 112 of the standby server 100 in which the abnormality has occurred.
[0078]
 FIG. 7 is a sequence chart showing an operation example in the case where an error occurs in all standby servers 100 in the cluster system 10 according to the second exemplary embodiment. The sequence chart shown in FIG. 7 is a sequence chart following the sequence chart shown in FIG. That is, the sequence chart shown in FIG. 7 is a sequence chart in a situation where a failure has already occurred in the standby server 100B. The operation of the cluster system 10 will be described below with reference to FIG. 7.
[0079]
 In step 401 (S401), a failure occurs in the standby server 100C, and the failure is detected by the monitoring process of the own server monitoring unit 113C of the standby server 100C.
 In step 402 (S402), the own server monitoring unit 113C notifies the cluster control unit 114C of the monitoring result (abnormality) performed in step 401.
 In step 403 (S403), the cluster control unit 114C notifies each of the other servers 100 of the monitoring result (abnormality) notified in step 402, and synchronizes the monitoring result.
[0080]
 Next, in step 404 (S404), the cluster control unit 114A confirms the synchronized monitoring result. Since the synchronized monitoring result indicates an abnormality in the standby server 100C, the cluster control unit 114A excludes the monitoring result by the other server monitoring unit 112C of the standby server 100C from the failover determination (exclusion flag). ).
[0081]
 Next, in step 405 (S405), since all standby servers 100 have become abnormal, the cluster control unit 114A switches to monitoring the business service by the own server monitoring unit 113A of the active server 100A. Therefore, the cluster control unit 114A requests the own server monitoring unit 113A of the active server 100A to start monitoring the business service. The cluster control unit 114A determines the necessity of failover based on the monitoring result of the own server monitoring unit 113A until one of the standby servers 100 becomes normal.
[0082]
 On the other hand, the cluster control unit 114C of the standby server 100C temporarily suspends the monitoring process of the other server monitoring unit 112C in step 406 (S406). The monitoring process of the other server monitoring unit 112C that is temporarily stopped here does not resume until the state of the standby server 100C returns to normal.
[0083]
 As described above, in the present embodiment, when there is no normal standby system server 100, it is possible to determine whether to perform failover based on the monitoring result of the local server monitoring unit 113 of the active system server 100. Therefore, even in a situation where the monitoring result of the other server monitoring unit 112 of the standby server 100 cannot be used, the necessity of performing failover can be determined.
[0084]
 Although the embodiments have been described above, the server apparatus having the configuration shown in FIG. 8 can also reliably determine whether the service is properly provided to the client. The server device 7 illustrated in FIG. 8 includes a service providing unit 8 (service providing unit) and a cluster control unit 9 (cluster control unit). The server device 7 constitutes a cluster system together with the standby server device.
[0085]
 The service providing unit 8 corresponds to the business service providing unit 120 of the above embodiment. The service providing unit 8 provides a predetermined service to the client device via the network.
 The cluster control unit 9 corresponds to the cluster control unit 114 of the above-described embodiment. The cluster control unit 9 acquires the monitoring result transmitted by the standby server device, and when the monitoring result indicates that the standby server device cannot normally access the predetermined service, performs a failover. Here, the standby server device is a device that takes over the provision of a predetermined service to the client device when a failover is performed, and accesses the predetermined service provided by the service providing unit 8 via the network. Then, monitor whether or not normal access is possible.
[0086]
 In this way, the server device 7 acquires the monitoring result of the service based on the access by the standby server device, and determines the implementation of failover. Therefore, the server device 7 can more reliably determine whether or not the service is properly provided to the client.
[0087]
 The present invention is not limited to the above-mentioned embodiments, but can be modified as appropriate without departing from the spirit of the present invention. For example, in the above embodiment, the HA cluster system is configured by the three servers 100, but the cluster system 10 only needs to have the active server 100 and the standby server 100, and the number of servers is It is optional. In addition, the server 100 is not limited to a one-way standby type, which is a cluster configuration in which the server 100 operates as any one of the active system and the standby system, and the server 100 is a cluster configuration in which the server 100 operates as the active and standby server. It is also possible to configure the cluster system 10 of the standby type.
[0088]
 Although the present invention has been described with reference to the exemplary embodiments, the present invention is not limited to the above. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the invention.
[0089]
 This application claims the priority on the basis of Japanese application Japanese Patent Application No. 2017-171129 for which it applied on September 6, 2017, and takes in those the indications of all here.
Explanation of symbols
[0090]
1, 10 Cluster system
2
, 100A Active server 3, 100B, 100C Standby server
4,
200, 300 Network 5, 9, 114, 114A, 114B, 114C Cluster control unit
6 Monitoring unit
7 Server device
8 Service providing unit
9 Cluster control unit
100 Server
110, 110A, 110B, 110C Clusterware
111, 111A, 111B, 111C Business service control unit
112, 112A, 112B, 112C Other server monitoring unit
113, 113A, 113B, 113C Own server monitoring unit
120, 120A , 120B, 120C Business service providing unit
151 Network interface
152 Memory
153 Processor
The scope of the claims
[Claim 1]
 When an
 abnormality occurs in the active server device that provides a predetermined service to the client device via the network and the active server device, the predetermined service is provided to the client device instead of the active server device. A standby system server device for providing
 ,
 wherein the standby system server device accesses the predetermined service provided by the active system server device via the network, and monitors whether or not normal access is possible. It has one of the monitoring means,
 the active server apparatus, when said active server apparatus is determined by the first monitoring means of the standby server apparatus and can not be accessed normally to the predetermined service provided , A
 cluster system having a cluster control means for performing failover .
[Claim 2]
 The cluster control means performs a failover when the first monitoring means of the plurality of standby server devices determines that the predetermined service provided by the active server device cannot be normally accessed.
 The cluster system according to claim 1.
[Claim 3]
 The cluster control unit is normally operated by the first monitoring unit of the standby system server devices in a number equal to or more than a predetermined ratio among the plurality of standby system server devices to normally perform the predetermined service provided by the active system server device.
 The cluster system according to claim 1 , wherein failover is carried out when it is determined that access to the cluster is impossible.
[Claim 4]
 The standby system server device further has a second monitoring means for monitoring the operating state of the standby system server device itself, and the
 cluster control means of the active server device is abnormal by the second monitoring means. Is detected, the first monitoring means of the standby system server devices of a predetermined ratio or more among the plurality of standby system server devices normally provide the predetermined service provided by the active system server device.
 The cluster system according to claim 3 , wherein failover is carried out when it is determined that access is impossible .
[Claim 5]
 The active server device further has third monitoring means for monitoring the operating state of the active server device itself, and the
 cluster control means of the active server device is abnormal by the second monitoring device.
 The cluster system according to claim 4 , wherein when there is no standby system server device for which a failure has not been detected, whether to execute failover is determined based on the monitoring result by the third monitoring means .
[Claim 6]
 6. The network is a public LAN, and is a network different from an interconnect LAN used for performing alive monitoring between the active server device and the standby server device
 . The cluster system according to item 1.
[Claim 7]
 An active server device provides a predetermined service to a client device via a network, and
 a standby server device that forms a cluster system together with the active server device provides the active server device via the network. If the
 active server cannot normally access the predetermined service provided by the active server, the standby server is monitored.
 A control method for a cluster system that performs failover if determined by .
[Claim 8]
 Service providing means for providing a predetermined service to a client device via a network, and
 a standby server for monitoring whether or not the predetermined service provided by the service providing means can be accessed normally via the network When the monitoring result transmitted by the device is acquired and the monitoring result indicates that the predetermined service cannot be normally accessed from the standby server device, the standby control device
 has a cluster control unit that
 executes failover , The system server device is a server device that is a device that takes over the provision of the predetermined service to the client device when a failover is performed
 .
[Claim 9]
 A predetermined service is provided to a client device via a network,
 and a monitoring result transmitted by a standby server device that monitors whether the predetermined service can be accessed and normally accessed via the network is acquired, When the monitoring result indicates that the standby service cannot normally access the predetermined service, failover is performed, and the
 standby server device performs the predetermined service when failover is performed.
 Control method, which is a device that takes over the provision of the client device to the client device .
[Claim 10]
 A service providing step of providing a predetermined service to a client device via a network, and
 monitoring whether the predetermined service provided by the processing of the service providing step is accessed via the network and can be normally accessed. When the monitoring result transmitted by the standby server device is acquired and the monitoring result indicates that the standby server device cannot normally access the predetermined service, a cluster control step for performing a failover is added
 to the computer.  A non-transitory computer-readable medium in which the
 standby server device
stores a program that is a device that takes over the provision of the predetermined service to the client device when a failover is performed .

Documents

Application Documents

# Name Date
1 202017010785-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [13-03-2020(online)].pdf 2020-03-13
2 202017010785-STATEMENT OF UNDERTAKING (FORM 3) [13-03-2020(online)].pdf 2020-03-13
3 202017010785-REQUEST FOR EXAMINATION (FORM-18) [13-03-2020(online)].pdf 2020-03-13
4 202017010785-PROOF OF RIGHT [13-03-2020(online)].pdf 2020-03-13
5 202017010785-PRIORITY DOCUMENTS [13-03-2020(online)].pdf 2020-03-13
6 202017010785-POWER OF AUTHORITY [13-03-2020(online)].pdf 2020-03-13
7 202017010785-NOTIFICATION OF INT. APPLN. NO. & FILING DATE (PCT-RO-105) [13-03-2020(online)].pdf 2020-03-13
8 202017010785-FORM 18 [13-03-2020(online)].pdf 2020-03-13
9 202017010785-FORM 1 [13-03-2020(online)].pdf 2020-03-13
10 202017010785-DRAWINGS [13-03-2020(online)].pdf 2020-03-13
11 202017010785-DECLARATION OF INVENTORSHIP (FORM 5) [13-03-2020(online)].pdf 2020-03-13
12 202017010785-COMPLETE SPECIFICATION [13-03-2020(online)].pdf 2020-03-13
13 202017010785-FORM 3 [28-08-2020(online)].pdf 2020-08-28
14 abstract.jpg 2021-10-19
15 202017010785.pdf 2021-10-19
16 202017010785-Power of Attorney-180320.pdf 2021-10-19
17 202017010785-OTHERS-180320.pdf 2021-10-19
18 202017010785-OTHERS-180320-1.pdf 2021-10-19
19 202017010785-OTHERS-180320-.pdf 2021-10-19
20 202017010785-Correspondence-180320.pdf 2021-10-19
21 202017010785-FER.pdf 2021-11-01
22 202017010785-OTHERS [31-01-2022(online)].pdf 2022-01-31
23 202017010785-FORM 3 [31-01-2022(online)].pdf 2022-01-31
24 202017010785-FER_SER_REPLY [31-01-2022(online)].pdf 2022-01-31
25 202017010785-COMPLETE SPECIFICATION [31-01-2022(online)].pdf 2022-01-31
26 202017010785-CLAIMS [31-01-2022(online)].pdf 2022-01-31
27 202017010785-Response to office action [24-04-2025(online)].pdf 2025-04-24
28 202017010785-US(14)-HearingNotice-(HearingDate-17-12-2025).pdf 2025-11-18

Search Strategy

1 202017010785E_25-10-2021.pdf